What SemiAnalysis actually measured on OpenAI's Jalapeño
SemiAnalysis verified OpenAI-supplied InferenceX runs in the lab. The authors say a Blackwell-only comparison is incomplete.
SemiAnalysis measured Jalapeño on OpenAI-supplied InferenceX runs in OpenAI's lab. The authors write that a comparison to Blackwell is “somewhat incomplete and unfair,” because the chip is still an HBM4 engineering sample while Vera Rubin systems that also use HBM4 are already shipping.
OpenAI Jalapeño: Better Than Nvidia Blackwell is the August 25, 2026 lab report by Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel. OpenAI invited the team after the Hot Chips announcement. Design work with Broadcom began in mid-2024 and reached tape-out in about 16 months. The chip is a generalized inference ASIC, not a specializer for OpenAI models.
What they measured
The published slice is single-token-prediction throughput on 8k-input, 1k-output prompts, without speculative decoding and without splitting prefill and decode across separate chip pools. SemiAnalysis verified those InferenceX runs in person. The numbers themselves come from OpenAI.
On tokens per megawatt, Jalapeño's single-token-prediction results beat Blackwell and the published Vera Rubin multi-token-prediction figures. DeepSeek R1 at concurrency 1 exceeded 700 tokens per second per user. The lab also ran Kimi K2.5 and GPT-OSS. GSM8k scores were on par with Nvidia chips.
Tokens per dollar land roughly even with Rubin on that same slice. Rubin already uses speculative decoding. SemiAnalysis says that technique can cut cost per token by about 3 to 5 times, so Jalapeño's cost picture can still move when it lands. OpenAI skipped prefill-decode disaggregation so the mix of input, cache, and output tokens can change without stranding a specialized pool.
What they did not measure
- The full InferenceX suite. The authors verified the runs they saw. They did not operate every benchmark themselves.
- AgentX. That long-context, multi-turn suite is the comparison they prefer for production-like cache behavior, and they have not seen Jalapeño results on it.
- A Blackwell-only scoreboard. They treat Rubin as the peer because both chips use HBM4, Rubin is shipping, and Jalapeño is still in engineering samples.
- Shipped production silicon. Production is scheduled to ramp through 2027.
How this sits next to the archive and the charts
The August 25 rough.day Tech & AI edition already selected the Hacker News story that led with Blackwell. This take keeps the authors' measurement and their limit. It does not replace that ranked item.
The Hraness reading note keeps the same source boundary: a general inference chip, tokens per megawatt as the scoreboard OpenAI is designing for, and the incomplete Blackwell comparison.
AI Charts is a checked coding-agent snapshot. That snapshot does not contain a Jalapeño row. Chip tokens-per-megawatt is a different measurement from coding-agent score, cost, and token use. This take does not invent a scoreboard row to close that gap.
Sources
- OpenAI Jalapeño: Better Than Nvidia BlackwellSemiAnalysis, 2026. The August 25, 2026 essay reports the lab visit, InferenceX verification, tokens-per-megawatt figures, total-cost comparison, and the authors' Blackwell-comparison limit.
- Hraness reading note: OpenAI Jalapeño: Better Than Nvidia BlackwellHraness, 2026. The Hraness reading note is a dated digest of the SemiAnalysis essay. It is a crawlable companion citation, not a substitute for the original.