A new system aimed at agents that run for hours
CoreWeave says NVIDIA Vera Rubin systems are now available for agentic AI workloads, with Cognition among the first customers running production work. The announcement focuses on the full cycle: training and evaluating agents, then serving the multi-step tasks those agents carry out.
NVIDIA reports up to 4.8 times the token throughput of GB200 NVL72 on selected SWE-2 coding-agent tasks. That is an “up to” result from a vendor test on particular tasks and systems, not a promise that every model or workload will run 4.8 times faster.
Why throughput matters for agent products
An agent may call tools, inspect results, revise a plan and continue for many turns. More output capacity can let infrastructure serve more concurrent work or shorten a long task, depending on the model, memory limits and system configuration. That is different from improving the model’s reasoning or the percentage of tasks it completes successfully.
CoreWeave’s Forge service is described as a way to train, evaluate and improve agents. The partnership is pitching infrastructure as part of the agent loop, where deployment feedback can inform later training and evaluation.
Read the benchmark as a narrow result
The public headline is a selected-task comparison. It does not disclose a universal benchmark across all software engineering tasks, and throughput is not the same metric as code quality. Buyers will need to compare actual task success, latency and total cost on their own workloads.
Still, the direction is clear: providers are building systems around long-running agents rather than single chat responses. The important next evidence will be customer results that report both throughput and completed-task quality under the same conditions.
Source published September 30, 2026. Coverage is based on the maker’s announcement and demonstration.
