On September 30, 2026, CoreWeave announced that its first Vera Rubin NVL72 capacity is running customer production workloads, with Cognition, the company behind the Devin coding agent, as the first customer. The headline number is that Cognition reports up to 4.8x the total token throughput of the previous-generation GB200 NVL72 on its SWE-2 inference workload. This piece looks at the CoreWeave Vera Rubin NVL72 launch itself: what was actually measured, under what conditions, how the number compares with other generational claims, and where the launch sits in the race among neoclouds to be first with Nvidia's newest rack. Earlier Metir coverage looked at the Nvidia Vera Rubin platform and its ramp and at Nebius validating a Rubin rack in Finland; this one is about the first customer-facing result.
NVIDIAWhat CoreWeave actually launched
According to Converge Digest, the announcement covers Nvidia Vera Rubin NVL72 racks on CoreWeave Cloud: 72 Rubin GPUs and 36 Vera CPUs per rack, 20.7 TB of HBM4 memory, roughly 1,400 TB/s of aggregate HBM4 bandwidth and 216 TB/s of NVLink 6 bandwidth, with ConnectX-9 SuperNICs and BlueField-4 DPUs. The same report says hundreds of Rubin GPUs are already deployed across multiple regions, and that Cognition's cluster was brought up in early September 2026.
Availability is described differently depending on the source. The press release is reported to claim general availability, while StorageReview notes that a CoreWeave blog post points to limited availability with gated allocation. Converge Digest also lists the launch as limited availability. For a buyer, that difference matters more than the throughput figure: a rack that exists is not the same as capacity a new customer can reserve this quarter.
The same event covered more than GPUs. Nvidia's blog on the launch says the Vera CPU is also coming to CoreWeave Cloud, at 128 CPUs and 11,264 cores per rack, aimed at the sandbox environments that coding agents run in. StorageReview adds that CoreWeave reports more than 3x faster agent sandbox startup on Vera than on an x86 CPU.
The Cognition workload
Cognition builds Devin, the autonomous software engineering agent covered in our post on its $1 billion run rate and $48 billion Series E. Agentic coding is a demanding inference pattern. Cognition's Silas Alberti, quoted by Converge Digest, described it as "an unforgiving workload that requires long contexts, high concurrency and rapid reasoning."
Per Nvidia's blog, the 4.8x figure was benchmarked on real-world software engineering tasks sampled from FrontierCode, with AI agents deployed to solve them. A second result, 3.8x higher output token throughput per GPU, applies to reinforcement learning, the training-side loop that agent labs use to improve models on tasks like these. Converge Digest says the inference comparison was made at matched interactivity, meaning the two racks were compared at the same responsiveness target rather than at whatever batch size flatters one of them.
The number is the customer's own, from early tests, and no one outside the two companies has checked it.
Mixed, on the 4.8x claim
The conditions behind 4.8x
Mixed frames the figure as a ceiling on one specific workload rather than a general improvement, measured by Cognition in its own early tests with no independent verification, against GB200 NVL72 racks Cognition already operates. StorageReview states that no model sizes, context lengths or batch settings were disclosed. Four conditions are worth keeping in view:
- "Up to." The phrase in Nvidia's wording is "up to a 4.8x increase," so it is a best case, not an average across tasks.
- Specific workload. The result is for SWE-2 inference. Other models, shorter contexts or lower concurrency may land elsewhere.
- Per GPU. Both GB200 NVL72 and Vera Rubin NVL72 are 72-GPU racks, so a per-GPU multiple should translate to a similar per-rack multiple (derived, not stated by the sources).
- Total versus output tokens. The inference figure counts total token throughput, which includes processed input, while the reinforcement-learning figure counts output tokens only. They are not the same unit, so 4.8x and 3.8x should not be read as one scale.
How it compares with other generational claims
Nvidia's own multiples use different metrics, so a straight comparison needs care. On its GB200 NVL72 page, Nvidia claims 30x for LLM inference versus H100, measured at a 50 millisecond token latency target with 32,768 input and 1,024 output tokens, and labels the figures projected performance subject to change. For Rubin, the January 2026 announcement claims up to a 10x reduction in inference token cost versus Blackwell and a 4x reduction in GPUs needed to train mixture-of-experts models.
Generational multiples, as claimed
Each bar is a different metric on a different workload. Green bars are NVIDIA statements, gray bars are Cognition early-test results. Compare them as claims, not as one scale.
30x: GB200 NVL72 vs H100, LLM inference at 50 ms token latency, 32,768 input / 1,024 output tokens (NVIDIA, projected). 10x: reduction in inference token cost (cost, not throughput). 4.8x and 3.8x: per GPU vs GB200 NVL72, no model size, context length or batch settings disclosed.
A derived check shows why these should not be merged. Cost per token is roughly price per GPU-hour divided by throughput per GPU. If Rubin delivered 4.8x the throughput on this workload, a 10x lower cost per token would require a Rubin GPU-hour priced at about 0.48 times a GB200 GPU-hour. Nvidia's 10x is described as a platform-level cost claim, and CoreWeave's June announcement framed it as up to 10x better inference per watt, so the claims are measuring different things. Cognition's figure is narrower, closer to what one buyer sees on one job, and it is smaller. That is the usual pattern when a customer replaces a vendor's projected best case with a measured workload, not evidence that either number is wrong.

The neocloud race to Rubin
CoreWeave's claim to be first rests on two dated milestones. On June 1, 2026, it said it was "the first AI cloud provider to bring up Vera Rubin," after system-level validation of the full rack-scale architecture. Roughly four months later, the September 30 launch moved from bring-up to a customer running production work. That is a company statement, and the sources do not independently confirm that no other provider completed a bring-up earlier.
Others are close behind. In its January announcement, Nvidia named AWS, Google Cloud, Microsoft and OCI among the first cloud providers to deploy Vera Rubin instances in 2026, and listed CoreWeave, Lambda, Nebius and Nscale as Nvidia Cloud Partners expected to deploy this year. Nebius reported validating its first Rubin NVL72 rack in Finland in late July, as covered in our earlier post. The distinction CoreWeave is now drawing is between a validated rack and a rack with a paying customer in production. Whether Lambda, Nebius or Nscale announce a comparable customer result is the next datapoint.
What to watch
- Independent verification. A third-party benchmark on comparable workloads would show how much of 4.8x is the hardware and how much is Cognition's tuning of its own stack.
- Real availability. The gap between a press-release availability claim and gated allocation will show up in lead times for other customers.
- Price per token. Throughput gains reach buyers only if hourly pricing does not rise to match.
- Second and third customers. Results from workloads other than agentic coding would test whether 4.8x is workload-specific.
- Backlog conversion. CoreWeave's contracted backlog is large, and our Q2 2026 earnings analysis covers whether new hardware speeds its conversion into revenue.
For teams that buy inference rather than racks, the practical lesson is that hardware multiples reach you through model providers and their pricing, which is one reason a model-agnostic workspace such as Metir AI lets you switch to whichever model gets cheaper or faster as new capacity comes online.
The takeaway
The verifiable facts are that CoreWeave has Vera Rubin NVL72 racks running a customer's production workloads, that Cognition reports up to 4.8x total token throughput per GPU on SWE-2 inference and 3.8x on reinforcement learning versus GB200 NVL72, and that the figures come from the customer's early tests without disclosed model, context or batch details. The result is a strong early signal for agentic coding and an unverified number for everything else. Treat it as the first measured datapoint of the Rubin generation, to be checked as more customers and independent tests arrive.
Sources:
- CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production with Cognition | Converge Digest
- CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production: Cognition Reports 4.8x Over GB200 | StorageReview
- From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI | NVIDIA Blog
- Cognition says Nvidia's Vera Rubin rack gave it up to 4.8x the token throughput of GB200 | Mixed
- CoreWeave Completes Industry-First Bring-Up and Validation of NVIDIA Vera Rubin NVL72 | CoreWeave
- NVIDIA Kicks Off the Next Generation of AI With Rubin | NVIDIA Newsroom
- NVIDIA GB200 NVL72 | NVIDIA
Image credits
Header and in-body image: Jensen Huang at the Nvidia keynote, CES 2025, Las Vegas, by Pronoia via Wikimedia Commons, released under CC0. It shows a consumer RTX Blackwell slide, not Vera Rubin hardware. Reviewed before use.