Nvidia spent Tuesday arguing that the number which matters for an AI data center has changed from raw throughput to tokens delivered per megawatt.
Ian Buck, the company’s vice president for hyperscale and high-performance computing, made the case at the AI Infra Summit in Santa Clara. The event drew more than 8,000 attendees this year, against 3,500 the year before.
His pitch was about a stack rather than a chip. Nvidia bundled Vera Rubin systems with Dynamo inference software, NeMo libraries and its networking line, listing NVLink for scale-up, Spectrum-X Ethernet, ConnectX SuperNICs, and BlueField storage and DPUs. The claim it all supports is validated agentic tokens per megawatt.
The specific number attached to that is DSX MaxLPS, a factory-wide power optimization layer that Nvidia says can lift output by up to 1.4 times more tokens per megawatt.
Electricity supply carried the rest of the message. Silicon Valley Power runs a flexible-load interconnection program that lets AI factories help balance the grid, and Emerald AI worked with Nvidia to demonstrate automated load reduction on that network. The system answered hundreds of demand signals while keeping workloads alive.
The framing lands as agentic workloads multiply and utilities grow slower about hooking up new capacity. Nvidia’s answer is to sell efficiency and grid cooperation as features of the platform rather than constraints placed on it.