On 25 August 2026, at the Hot Chips conference, OpenAI published the first measured results for Jalapeño, its custom inference chip co-developed with Broadcom. The headline claims are specific, drawn from a public benchmark, and genuinely impressive: 1.5 to 1.9 times more AI work per watt at peak throughput, and 1.7 to 3.6 times lower end-to-end latency, across three open-weight models.
Those two ranges will be quoted everywhere. This article is about the sentence printed directly underneath them, which will be quoted nowhere.
“To compare the systems consistently, we normalized the results using each accelerator’s published chip power rating. Jalapeño is rated at 700 watts, although its measured sustained power remained at or below 550 watts on the workloads tested.”
The denominator is a spec sheet
Sit with what that sentence concedes. The efficiency claim is expressed per watt. The watts in the denominator are nameplate thermal design power — 700W for Jalapeño, 1,200W for the GB200, 1,400W for the GB300 — not measured draw. In the same breath, OpenAI discloses that Jalapeño’s actual sustained consumption stayed at or below 550W, meaning its own nameplate overstates its real draw by roughly a fifth.
The obvious reading is that this makes the comparison conservative: OpenAI penalised its own chip by using the larger number. Several outlets have already reached that conclusion and framed the real efficiency edge as larger than published. That reading is not wrong so much as unsupported, and the reason is the thing worth understanding.
OpenAI publishes a measured figure for Jalapeño. It publishes no equivalent measured figure for the comparison systems. Datacentre GPUs also routinely draw less than nameplate on real inference workloads — how much less depends on the model, the batch size, the memory access pattern and the cooling envelope. Without measured draw on both sides, the direction of the error is unknown. If the comparison systems happened to run further below their nameplate than Jalapeño ran below its own, the normalisation flatters Jalapeño rather than penalising it. Nobody outside these two companies can currently say which.
This is not an accusation of dishonesty. OpenAI volunteered a caveat that most vendor benchmarks omit entirely, and it named the benchmark, the models and the comparison hardware. The point is narrower and more uncomfortable: ‘performance per watt’ is currently an unfalsifiable unit, because whether it means measured power or a number on a datasheet changes the answer, and the industry has not agreed which it means.
That generalises far past this chip. Every efficiency claim in AI infrastructure right now — vendor decks, sustainability disclosures, total-cost-of-ownership models, datacentre siting arguments — rests on the same unexamined denominator. If you are shown a per-watt figure in the next month, there is exactly one question that determines whether it means anything: measured or nameplate? Ask it about both sides of the comparison.
How a 1.7× becomes a 104×
The appendix of OpenAI’s post contains figures like ≈53.7× and ≈104.3×. They are arithmetically correct. They will escape into headlines, and when they do they will be badly misunderstood.
The label on those numbers is “more throughput at previous TBT” — throughput measured at one specific matched latency point, where TBT is time between tokens. The chosen point is the comparison system’s previous best time-between-tokens, which is precisely where that system falls off a performance cliff. Ratios taken at a cliff edge are enormous by construction. They describe the steepness of one system’s degradation curve at one operating point, not a general capability gap.
Here is the same table read honestly, as OpenAI reports it:
- GPT-OSS 120B vs GB200 (1,200W): ≈1.9× peak mixed TPS/kW — 85,448 against 44,960 — and ≈1.7× lower end-to-end latency, 1.03s against 1.80s. The appendix ratio at matched TBT: ≈53.7×.
- DeepSeek R1 670B vs GB300 (1,400W): ≈1.7× — 19,641 against 11,781 — and ≈3.6× lower latency, 1.65s against 5.99s. Matched-TBT ratio: ≈104.3×.
- Kimi K2.5 1T vs GB300: ≈1.5× — 18,195 against 11,862 — and ≈3.4× lower latency, 1.56s against 5.31s. Matched-TBT ratio: ≈56.1×.
The honest headline number from that table is 1.5 to 1.9 times. It is a good number. It does not need help. The benchmark itself — InferenceX, published by SemiAnalysis — is public, which is a real point in OpenAI’s favour; but the runs are OpenAI’s own, on OpenAI’s unreleased hardware, and have not been independently reproduced. The correct verb is ‘OpenAI reports,’ not ‘Jalapeño is.’
What is actually new here
Strip out the benchmark theatre and two claims remain that deserve more attention than the per-watt ratio.
The first is schedule. OpenAI reports nine months from initial design to tapeout for a reticle-scale inference ASIC. For custom silicon of this class that is a remarkable cycle time, and it is the kind of claim that is hard to fake — the chip either exists on that timeline or it does not. OpenAI says its own models were used to explore implementations and optimise arithmetic circuits during that process.
The second is more interesting and more carefully hedged. For selected GPT-OSS attention and mixture-of-experts blocks, OpenAI says AI-generated kernel implementations ran 1.5 to 1.8 times faster than the existing human-expert-written ones. The company immediately adds its own caveat, and it belongs in every retelling: “Those figures apply to the selected blocks, not the full model.” Selected blocks are chosen; a speedup on hand-picked kernels is not a speedup on a model. It is still a meaningful data point about where machine-generated low-level code now sits relative to specialists. OpenAI also says three open-weight models outside the original production plan were brought to high performance within two months using Codex with GPT-Astra.
This is not a Nvidia displacement story, and OpenAI says so
It is worth being blunt about the narrative this will be recruited into, because OpenAI closes that door itself in the same post: “We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads.”
The practical status supports that. Jalapeño is inference-only. It is first-generation. It is built by Broadcom. It is not deployed — OpenAI says deployment inside its own infrastructure begins by the end of the year, and that it is still “continuing production qualification.” Gen 2 is described as deep in development and Gen 3 as taking shape, which is what you would expect and is not evidence about Gen 1.
The real story is vertical integration and design velocity: a model company that can now specify, co-design and tape out inference silicon on a nine-month cycle has changed its own cost structure and its negotiating position, without changing what it buys this year. That is a slower, less dramatic and considerably more consequential development than a benchmark table.
And when the Jalapeño numbers land in a vendor deck in front of you — they will — the question to ask is still the first one. Measured, or nameplate?