CoreWeave offers Nvidia's Vera Rubin NVL72, with Cognition as first customer
Cognition says its SWE-2 inference workloads saw up to a 4.8x increase in total token throughput on the new systems.
CoreWeave has made Nvidia's Vera Rubin NVL72 generally available on its cloud, about two months after it said it had built and was operating what it described as the first fully working rack of the system, and it has landed a customer already running production work on the hardware. The company announced the availability at its Fully Connected conference in San Francisco, and Cognition, the first adopter, began using the Vera Rubin systems in early September.
The figure that will travel furthest is a performance number from Silas Alberti, an SVP of research and a member of Cognition's founding team, who says the company has seen up to a 4.8x increase in total token throughput for its SWE-2 inference workloads on the new generation. It is one customer's measurement on its own workloads rather than a neutral benchmark, and the sort of claim that does the selling for a next-generation rack before broad deployment data exists. Chen Goldberg, CoreWeave's executive vice president of product and engineering, framed the fast bring-up as the payoff from years of engineering the platform across GPU generations, and said the aim is to make compute, networking, and software work as a single system so customers can build complex agents without absorbing the infrastructure complexity themselves.
The Rubin platform spans six chips: the NVLink 6 switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU, and the Spectrum-6 Ethernet switch, alongside the Vera CPU and the Rubin GPU. Vera succeeds Nvidia's Grace CPU and Rubin succeeds its Blackwell GPUs; Nvidia claims Rubin will deliver 5x inference and 3.5x training performance against Blackwell. The NVL72 rack itself pairs 36 CPUs with 72 GPUs, is 100 percent liquid-cooled, and uses cable-free modular tray designs that cut installation time from two hours to five minutes.
For neoclouds, speed to Nvidia's newest silicon is the differentiator worth advertising, and the two-month gap between the first fully working rack and general availability is itself part of the pitch. A customer sizing a large inference job is choosing, in part, how soon it can run on the current generation; CoreWeave's prior claim that its rack was the first fully working one carried the same implied comparison, and now the company has a customer's throughput figure to hang on it.
A five-minute rack does not shorten the queue
Those design details aim at the people who own the buildings and compress the wrong part of the schedule: when Schneider standardized its factory-built 2.5MW power modules, the point still held that squeezing the most controllable step leaves the interconnection queue exactly where it found it. A rack that installs in five minutes is still a rack waiting on a feeder, and full liquid cooling is a density tell—the more heat a rack moves into water rather than air, the more the site's power and cooling design, rather than the chip, sets when the capacity starts earning. That likely means a facility built around air handling needs new cooling distribution before these racks go in, and Siemens Grid Software's planning guidance puts AI racks at 230 kilowatts and recommends scenario-based planning for large loads. Rack-scale systems do not change that arithmetic so much as move it to the front of the project.
Vera CPU becomes a bare-metal product
The quieter item in the packet is, for CoreWeave's business mix, more consequential: the company plans to offer Nvidia's Vera CPU as a standalone, bare-metal product, a departure for a neocloud that has traditionally sold access to GPUs. Corey Sanders, CoreWeave's SVP of product, calls the effort early: some customers are expected to start testing in the coming weeks, and a few already have early access. Vera is designed for AI agents and the workloads that arrive with them, and running it bare-metal on CoreWeave puts a CPU product beside the GPU hours the company has built its order book around.
For the capital stack behind racks like these, the name matters more than the SKU: data center capital now prices the anchor rather than the megawatt, with hyperscaler- or AI-lab-anchored assets drawing infrastructure treatment while everything else fights for terms. A cloud contract is not a lease, and the coverage does not include a contract value, so nothing here settles a financing. The announcement instead supplies a next-generation rack going live with a named AI lab already running inference workloads on it, the sort of reference an anchor claim needs.
Two numbers will carry into the next round of capacity decisions: Nvidia's claimed 5x inference and 3.5x training improvement over Blackwell, and Cognition's 4.8x on its own inference workloads. Watch the Vera CPU, where Sanders says testing should begin in the coming weeks; the first named customers on that offering will show whether CoreWeave's move past GPU hours finds a market of its own.
Save this analysis and keep the funds you follow together in My Desk.
Sign in to save articles or follow funds.