OpenAI Designed Jalapeño Chip Using Its Own LLMs — SkimNews
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI unveiled Jalapeño on 25 August — its debut AI accelerator delivering 13.4 petaflops of 4-bit compute with 232 GB of advanced memory linked at 15.4 TB/s, with company benchmarks claiming up to 3.6x lower end-to-end latency than Nvidia's GB300 at lower power.
- The Jalapeño timeline spanned under 20 months from first architecture concept to first silicon, with only 9 months separating first RTL from tape-out, executed by a team averaging fewer than 100 people — Broadcom handled physical design from the gates onward while OpenAI led system design.
- OpenAI's LLMs drove gains through high-level synthesis tools like XLS (originally developed at Google), and on DeepSeek's multi-head latent attention kernel benchmark, software optimization performance climbed from 0.31% to 88.94% of theoretical ceiling in roughly 40 hours after first silicon returned in May.
- The project started with assistance from OpenAI's o3 model and by completion ran on precursors to GPT-6 Astra — a model that, unlike o3, can work directly in Verilog without needing XLS's translation layer and is close to operating proprietary design tools on its own.
- Richard Ho and Chris Leary reported at IEEE Hot Chips 2026 that AI-guided physical design achieved a 10% area reduction for matrix multiplication units versus an optimized human baseline; Broadcom's team used OpenAI's commercial models since it lacked access to the internal fine-tuned LLMs.
Why it matters: Jalapeño's claimed 3.6x latency edge over Nvidia's GB300 — if validated in production inference — puts established accelerator vendors on notice that custom AI-designed silicon has arrived. The 20-month concept-to-silicon timeline, executed with roughly 100 designers and Broadcom's physical-design help, demonstrates a compression of custom-chip schedules the entire semiconductor industry is now watching.
Ask SkimNews



