OpenAI's Jalapeño Chip Beats Nvidia on Inference Tests

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI introduced Jalapeño, an ASIC chip built with Broadcom for AI inference and first revealed in June, with hardware VP Richard Ho calling it the 'best of both worlds' on latency and throughput
- Jalapeño delivered 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 and GB300 superchips on the InferenceX benchmark across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T models
- The chip achieved 1.7 to 3.6 times lower end-to-end latency than the Nvidia systems across the same three models, which Ho said translates to faster responses and more reliable agent access
- OpenAI plans small-volume Jalapeño deployment by end of 2025, with volume ramping through 2027, though the company did not disclose specific deployment figures for next year
- Ho explicitly said OpenAI does not expect to replace its chip lineup with Jalapeño, calling Nvidia a 'very good partner' and confirming second and third generations of the chip are in development
Why it matters: Jalapeño's benchmark numbers give OpenAI a credible case to in-house more of its inference stack, but Ho's insistence that Nvidia remains a 'very good partner' signals Jalapeño is additive rather than a replacement — meaning OpenAI will run a mixed chip fleet through at least 2027 as it scales its own silicon from small volumes.
Ask SkimNews


