What SemiAnalysis measured on OpenAI's Jalapeño
SemiAnalysis verified OpenAI-supplied InferenceX runs in the lab. The authors say a Blackwell-only comparison is incomplete.

SemiAnalysis measured Jalapeño on OpenAI-supplied InferenceX runs in OpenAI's lab. The authors write that a comparison to Blackwell is “somewhat incomplete and unfair,” because the chip is still an HBM4 engineering sample while Vera Rubin systems that also use HBM4 are already shipping.
OpenAI Jalapeño: Better Than Nvidia Blackwell is the August 25, 2026 lab report by Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel. OpenAI invited the team after the Hot Chips announcement. Design work with Broadcom began in mid-2024 and reached tape-out in about 16 months. The chip is a generalized inference ASIC, not a specializer for OpenAI models.
What they measured
The published slice is single-token-prediction throughput on 8k-input, 1k-output prompts, without speculative decoding and without splitting prefill and decode across separate chip pools. SemiAnalysis verified those InferenceX runs in person. The numbers themselves come from OpenAI.
On tokens per megawatt, Jalapeño's single-token-prediction results beat Blackwell and the published Vera Rubin multi-token-prediction figures. DeepSeek R1 at concurrency 1 exceeded 700 tokens per second per user. The lab also ran Kimi K2.5 and GPT-OSS. GSM8k scores were on par with Nvidia chips.
Tokens per dollar land roughly even with Rubin on that same slice. Rubin already uses speculative decoding. SemiAnalysis says that technique can cut cost per token by about 3 to 5 times, so Jalapeño's cost picture can still move when it lands. OpenAI skipped prefill-decode disaggregation so the mix of input, cache, and output tokens can change without stranding a specialized pool.
What they did not measure
- The full InferenceX suite. The authors verified the runs they saw. They did not operate every benchmark themselves.
- AgentX. That long-context, multi-turn suite is the comparison they prefer for production-like cache behavior, and they have not seen Jalapeño results on it.
- A Blackwell-only scoreboard. They treat Rubin as the peer because both chips use HBM4, Rubin is shipping, and Jalapeño is still in engineering samples.
- Shipped production silicon. Production is scheduled to ramp through 2027.
What the result supports
The report supports a narrow conclusion: on the published test slice, the engineering sample delivered strong single-token-prediction throughput per megawatt. It does not establish production cost, AgentX performance, or a complete comparison with Nvidia's current generation. Those questions require results from shipped silicon and like-for-like inference methods.
Sources
- OpenAI Jalapeño: Better Than Nvidia BlackwellSemiAnalysis, 2026. The August 25, 2026 essay reports the lab visit, InferenceX verification, tokens-per-megawatt figures, total-cost comparison, and the authors' Blackwell-comparison limit.