AI is now capable of developing its own inference hardware — SkimNews
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- openTPU is an open-source AI accelerator whose hardware was designed by AI agents using an 'auto-arch-tournament' methodology, with the full stack — SystemVerilog, ISA, bit-exact simulator, kernel language and compiler, and host software — living in one monorepo under Apache 2.0.
- Build B (deploy_fused133c_79c5707a, in production since 2026-10-01) decodes LFM2-2.6B, SmolLM3-3B and Phi-4-mini 8-9% faster than the prior image (Gemma 4 E2B by 10%), at 91-94% of DDR3 peak versus 82-87% before, on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels).
- The new image achieves decode within 2.3% of the previous production image (se-cand3, Xilinx MIG, two-column matrix unit) in every configuration, while prefill runs 1.3x (Qwen3.5) to 2.0x (LFM2 4-bit) faster, and the on-card CPU calibrates both DDR3 channels in 12 seconds with no host involvement.
- Mixture-of-experts models larger than the card's 4 GiB run with experts streamed from host storage — LFM2.5-8B-A1B at 10.6 tok/s and Qwen3.5-35B-A3B at 3.95 tok/s (153 MB streamed per token at 1.41 GB/s over PCIe), with both matching the simulator bit for bit.
- 4-bit weights (FP4 with two-level block scales, 4.25 bits per weight) cut bytes per token by about a third and raise decode speed by 40% on Qwen3.5 to 45% on Qwen3 and LFM2, at a measurable perplexity cost documented per model.
- The design runs ten modern models whose on-card outputs match the bit-exact simulator token for token, including Hugging Face's greedy outputs on three prompts for Gemma 4 E2B.
Why it matters: openTPU shows AI agents can design functional inference silicon that runs production-scale LLMs bit-exact against a software simulator on a mid-range Kintex-7 xc7k480t FPGA. The 91-94% DRAM efficiency at 133.33 MHz — and prefill gains of 1.3-2.0x over a Xilinx MIG reference build — demonstrate that AI-designed accelerators can compete with hand-tuned vendor IP on real workloads, while the Apache 2.0 monorepo puts a complete hardware-to-runtime reference design in any researcher's hands.
Ask SkimNews
