AI

OpenAI Unveils Jalapeno Custom Inference Chip, Claims 1.9x Performance Per Watt Over Nvidia

OpenAI disclosed measured results from its first-party Jalapeno ASIC on August 26, saying the 700W chip delivers 1.5-1.9x more AI work per watt and 1.7-3.6x lower latency than competing systems across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T.

S
By Sarah Chen Senior AI Reporter
August 26, 2026 / 6 min read

OpenAI on August 26 disclosed the first measured results for Jalapeno, its in-house inference ASIC co-developed with Broadcom, claiming the 700W chip delivers 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than competing systems at peak throughput. Sam Altman announced the results on X with the line "we made a chip, and it is fast," and the company posted a paper on its website describing the architecture and benchmark methodology. Tested on the public SemiAnalysis InferenceX suite, the chip sustained power at or below 550W on the workloads evaluated.

Benchmarks and Numbers

OpenAI said Jalapeno was evaluated across three models — GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T — and three operating regimes ranging from high-throughput batch serving to highly interactive low-latency use. Across all three models, Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, OpenAI said, Jalapeno delivered 2.1 to 4.1 times higher performance. On the largest public model tested, Kimi K2.5 1T, the chip delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system.

How the Chip Was Built

Jalapeno is rated at 700W with measured sustained power at or below 550W on tested workloads, and OpenAI said the design was built to minimize data movement and communication delays, with model state including the KV cache explicitly placed and kept local. The company said the chip, its memory, network, software, and rack-scale system were developed together, allowing OpenAI to "co-design" the stack. OpenAI highlighted the role of AI in developing the chip itself, saying AI tooling helped the team move from initial design to tapeout in nine months and brought three open-weight models not part of the original production plan to high performance within two months.

What Comes Next

OpenAI said it plans to begin deploying Jalapeno within its compute infrastructure by the end of the year, calling it "the first generation of a multigenerational roadmap" with Gen 2 "deep in development" and Gen 3 "taking shape." The company emphasized that it will continue to deploy accelerators from Nvidia and other partners for training and inference, while preparing Jalapeno for large-scale operation and validating performance across additional models. OpenAI did not benchmark Jalapeno against Nvidia's Vera Rubin platform, which is slated to power the first gigawatt of Nvidia systems OpenAI agreed to deploy in the second half of 2026, Tom's Hardware noted.

What to Watch Through Year-End

Three checkpoints follow. The first Jalapeno deployments inside OpenAI's own production stack, expected by year-end 2026, will be the first real test of whether the published InferenceX numbers hold up under live traffic. Broadcom's role as co-developer — and the broader implication that hyperscalers now see custom inference silicon as essential — will set the tone for next year's ASIC roadmap announcements from Anthropic, Meta, and Google. The Gen 2 and Gen 3 tapeout schedules, which OpenAI described only as "deep in development" and "taking shape," will determine whether OpenAI's silicon cadence can match Nvidia's roughly annual Vera Rubin-style refresh cycle.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.