Jalapeño Gives Every Nvidia Customer a New Sentence to Say
The benchmark numbers have caveats. The negotiating leverage does not. The pattern points to Nvidia losing unilateral control over AI chip roadmaps before a single Jalapeño ships at scale.

TITLE: Jalapeño Gives Every Nvidia Customer a New Sentence to Say
At the Hot Chips conference yesterday, OpenAI put real numbers behind Jalapeño. Not slides. Numbers on silicon, verified partly in-person by a third party.
OpenAI says Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than commercially available systems. For interactive workloads, performance is 2.1x to 4.1x higher. The models tested were GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.
The verdict: the largest Nvidia customer just proved it can build competitive silicon.
What the Numbers Actually Say
Read the caveats before you forward the press release. Jalapeño remains at the engineering-sample stage, with deployment starting by end of 2026 and larger rollout in 2027. Engineering samples beating production Blackwell is routine; margins often reverse at volume.
SemiAnalysis points out that the fairer comparison is Nvidia's Vera Rubin platform, since both use HBM4 memory. Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, though Nvidia's accelerator uses multi-token prediction optimization that Jalapeño has not yet adopted. On total cost of ownership per token, the two come out roughly even.
TCO equivalence is the real number. If cost-per-token is the same, Nvidia wins on installed base, software maturity, and fifteen years of CUDA tooling. That Jalapeño reaches parity at all is the story.
All numbers come from OpenAI. SemiAnalysis verified the InferenceX runs in person but did not run the full benchmark suite. Their preferred comparison framework is AgentX, which tests long-context and multi-turn characteristics. That gap has not been closed in public yet.
Nvidia has not let SemiAnalesis test and release Rubin benchmarks the way OpenAI has, which suggests their chip software is still immature. Two chips at roughly similar stages of bring-up. One is letting independent testers in. The other is not.

The Leverage Is Already Real
The consensus treats this as a product story. The pattern points to procurement.
The June unveiling was the payoff of a deal OpenAI and Broadcom announced in October 2025 — a collaboration to deploy 10 gigawatts of OpenAI-designed AI accelerators, spanning both 3-nanometer chips and a future 2-nanometer generation. This is a multi-year platform bet with a named partner and a foundry roadmap.
Translation: when OpenAI says it will "continue using accelerators from Nvidia and other partners," it means Nvidia now negotiates alongside other partners, not above them.
OpenAI taped out Jalapeño in November 2025. Within nine months of tape-out and three months of bring-up on actual silicon, it delivered competitive results. OpenAI used AI to help design, architect, and optimize the chip, cutting the cycle from inception to tape-out to nine months — a timeline that conventional ASICs stretch to two or three years. That speed compresses the risk calculus for every other lab watching.
You do not need Jalapeño to ship at full volume to feel this in a capex conversation with Nvidia next week. You need it as a credible counteroffer. It is.
The timing — benchmarks dropped during annual capex planning season, two months before Nvidia's next earnings cycle — reads as rewriting terms before contracts renew. That is not accident. That is strategy.
What to Watch
Watch whether Nvidia's next earnings call contains new language around customer concentration or inference chip pricing. Watch whether Meta, Google, or Microsoft announce accelerated custom silicon programs within six months. And watch whether Nvidia cuts inference pricing quietly before Jalapeño scales, which would mean the chip already won without shipping a single rack.
- OpenAI's first custom chip 'Jalapeño' reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks
- OpenAI Jalapeño: Better Than Nvidia Blackwell
- OpenAI's New Jalapeno Chip Beats NVIDIA's Blackwell On Some Parameters, Company Says
- OpenAI and Broadcom Unveil LLM-Optimized Intelligence Processor
- OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast
- Wall Street Lunch: OpenAI's JalapeñO Tops Nvidia's Blackwell In Some Inference Tests | Seeking Alpha
- OpenAI's Jalapeño Chip Beats Nvidia Blackwell on Key AI Benchmarks
- OpenAI’s Jalapeño AI Chip Outperforms Nvidia Blackwell in Early Tests | citybiz
- OpenAI Unveils Jalapeño Chip Benchmarks, Claiming Big Efficiency Gains Over Nvidia — BigGo Finance
- Inference Chip Wars: OpenAI Jalapeño's Amazing Speed Win
- OpenAI and Broadcom unveil LLM-optimized inference chip | OpenAI
- Broadcom and OpenAI unveil custom-built Jalapeño inference processor — OpenAI's first chip is a massive reticle-sized ASIC built in an ultra-fast nine-month development cycle | Tom's Hardware
- OpenAI Jalapeno Chip Explained: Specs, Cost, Release Date
- Deep Dive into OpenAI's First Custom AI Chip Jalapeño: 9-Month Tape-out, 50% Inference Cost Reduction, AI Designing AI Hardware | HappyRock
- OpenAI Delivers Samples of First Custom Chip Jalapeno, Outperforming Nvidia's Existing Products in Some Tests