OpenAI's Jalapeño Beats Blackwell on AI Work per Watt
The first published benchmarks for OpenAI's Broadcom-built inference chip show 1.5× to 1.9× more work per watt than the best shipping systems, at half the package power of an Nvidia GB300. The caveats are real — engineering samples, 8k-context tests, no agentic workloads — but a first-generation accelerator is not supposed to land this close to Rubin.
OpenAI's first in-house accelerator has its first published benchmark numbers, and they are better than a debut chip has any business being. Jalapeño, the inference-only part OpenAI designed with Broadcom, delivered 1.5× to 1.9× more AI work per watt at peak throughput than the best commercially available systems, according to results run on SemiAnalysis's InferenceX suite. It drew 700 watts per package doing it, against 1,400 watts for Nvidia's GB300 — an efficiency gap that exists before any performance figure is counted.
Latency told a similar story. Jalapeño came in with 1.7× to 3.6× lower end-to-end latency than the systems it was measured against, and 2.1× to 4.1× higher performance on interactive workloads — the single-user, answer-now pattern that makes a chatbot feel fast rather than merely productive. On DeepSeek R1 the chip cleared 700 tokens per second for a single concurrent request; on GPT-OSS 120B it reached roughly 1,400 tokens per second per user. The three models tested were GPT-OSS 120B, DeepSeek R1 at 670B, and Moonshot's Kimi K2.5 at a trillion parameters.
The silicon underneath is unglamorous in the way good infrastructure usually is. The compute die is a single reticle-sized part on TSMC's N3P node, paired with an N3E I/O chiplet carrying 32 lanes of 800G SerDes. Memory is Samsung HBM4 running at 10Gbps per pin — slightly quicker than the 9.6Gbps Nvidia uses in Rubin — for 15.4TB/s of bandwidth per package, feeding a quoted 13.4 PFLOPS of MXFP4. Racks hold 128 accelerators across sixteen eight-chip trays, switched by Broadcom Tomahawk 6 silicon, at about 160kW per rack; the fabric scales to 2,048 chips across sixteen racks.
The caveats deserve as much attention as the multipliers, and SemiAnalysis flagged most of them. Every number came from OpenAI, though the analysts say they watched the InferenceX runs in person. Only engineering samples exist, while Nvidia's Rubin is already shipping to customers — comparing a lab part to a product in the field flatters the lab part. The tests were confined to 8,000 tokens of context and 1,000 of output, with no results on the longer-context, multi-turn workloads that actually characterise agentic production traffic. And Jalapeño posted these figures without multi-token prediction or speculative decoding, optimisations its competitors were using, which cuts both ways: the headroom is real, but so is the fact that the comparison systems were not running naked.
Against Vera Rubin specifically — the fairer matchup, since both parts use HBM4 — the picture narrows to something closer to parity on total cost of ownership, with Jalapeño still squeezing out more output tokens per megawatt. That is a more modest claim than beating Blackwell, and a more interesting one. OpenAI first unveiled Jalapeño in June, roughly nine months after the design work began in mid-2024 and following fabrication last November.
What the chip cannot do is train anything; it is an inference part, and OpenAI has been careful not to frame it as an exit from anyone's supply chain. CFO Sarah Friar has positioned it as complementary to the company's existing arrangements with Nvidia, AMD, AWS, Cerebras and CoreWeave rather than a substitute for them. Wall Street is reading it less politely: CNBC framed the results as a fresh threat to Nvidia's margins, on the theory that the largest buyers of inference capacity building credible silicon of their own changes the negotiation even when they keep signing the contracts. The number worth watching is not the perf-per-watt multiplier but whether Jalapeño holds up on long-context agentic traffic — the workload OpenAI actually runs, and the one nobody has benchmarked yet.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.