Skip to content
Rough Day
Theme
Appearance

What SemiAnalysis measured on OpenAI's Jalapeño

SemiAnalysis verified OpenAI-supplied InferenceX runs in the lab. The authors say a Blackwell-only comparison is incomplete.

Published , updated

Two-column diagram of what SemiAnalysis verified for OpenAI Jalapeño and what that published test slice did not establish.
The SemiAnalysis report verifies one InferenceX lab slice; it does not establish AgentX, the full suite, shipped silicon, or a complete Blackwell-only comparison.Code-native editorial diagram by Rough Day. Sources: SemiAnalysis.

SemiAnalysis measured Jalapeño on OpenAI-supplied InferenceX runs in OpenAI's lab. The authors write that a comparison to Blackwell is “somewhat incomplete and unfair,” because the chip is still an HBM4 engineering sample while Vera Rubin systems that also use HBM4 are already shipping.

OpenAI Jalapeño: Better Than Nvidia Blackwell is the August 25, 2026 lab report by Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel. OpenAI invited the team after the Hot Chips announcement. Design work with Broadcom began in mid-2024 and reached tape-out in about 16 months. The chip is a generalized inference ASIC, not a specializer for OpenAI models.

What they measured

The published slice is single-token-prediction throughput on 8k-input, 1k-output prompts, without speculative decoding and without splitting prefill and decode across separate chip pools. SemiAnalysis verified those InferenceX runs in person. The numbers themselves come from OpenAI.

On tokens per megawatt, Jalapeño's single-token-prediction results beat Blackwell and the published Vera Rubin multi-token-prediction figures. DeepSeek R1 at concurrency 1 exceeded 700 tokens per second per user. The lab also ran Kimi K2.5 and GPT-OSS. GSM8k scores were on par with Nvidia chips.

Tokens per dollar land roughly even with Rubin on that same slice. Rubin already uses speculative decoding. SemiAnalysis says that technique can cut cost per token by about 3 to 5 times, so Jalapeño's cost picture can still move when it lands. OpenAI skipped prefill-decode disaggregation so the mix of input, cache, and output tokens can change without stranding a specialized pool.

What they did not measure

What the result supports

The report supports a narrow conclusion: on the published test slice, the engineering sample delivered strong single-token-prediction throughput per megawatt. It does not establish production cost, AgentX performance, or a complete comparison with Nvidia's current generation. Those questions require results from shipped silicon and like-for-like inference methods.

Sources

  1. OpenAI Jalapeño: Better Than Nvidia BlackwellSemiAnalysis, 2026. The August 25, 2026 essay reports the lab visit, InferenceX verification, tokens-per-megawatt figures, total-cost comparison, and the authors' Blackwell-comparison limit.