OpenAI finally showed what its in-house inference silicon can do.
At the Hot Chips conference this week the company released the first benchmark results for Jalapeno, the custom chip it built with Broadcom. On SemiAnalysis’ InferenceX benchmark, the silicon delivered more tokens per user and higher throughput per kilowatt than an Nvidia Blackwell system.
Hardware chief Richard Ho described the gap as “a very, very significant performance advance over state of the art,” with more AI work served per unit of power and faster responses.
Announced last October, Jalapeno was designed with help from OpenAI’s own models and is meant to become a multigenerational platform spanning chips, memory, and models. The design attacks the usual inference bottlenecks: prefill and communication. Model state, including the KV cache, can stay local while the system reconfigures compute, memory, and networking per phase.
Ho expects very small deployments by the end of 2026, with real scale in 2027. The benchmark baseline is today’s Blackwell, so rivals will not stand still in the meantime.