NVIDIA is making the efficiency case for its next platform. Vera Rubin NVL72 racks, the company says, deliver up to 30x the inference throughput per megawatt of its current GB300 NVL72 on agentic workloads.
The measurements come from the SemiAnalysis AgentX workload, which replays recorded agentic coding sessions with real context growth, tool calls and sub-agent spawning. Agents are the token hogs of the moment, consuming about 15x more than a simple chat request as they bounce through tools, sub-agents and long context windows.
NVIDIA also claims the platform cuts cost per million tokens by up to 35x versus GB300 NVL72, helped by NVFP4 4-bit quantization, fifth-generation Tensor Cores and a third-generation Transformer Engine. Its power management software can fit up to 40% more GPUs into a given megawatt envelope.
Seven chips make up the full platform: the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC, joined by the NVL72 scale-up domain and sixth-generation NVLink at 10x bandwidth. Full production is underway and scaling across the ecosystem, NVIDIA says.
SemiAnalysis has yet to review the early figures, and the results exclude Vera CPU performance for tool calling.