The CPU is back: Rethinking the CPU-GPU split for LLM inference

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Intel reports CPU-to-GPU ratios in AI data centers are shifting from ~1:8 in training to ~1:1 in agentic workloads, with some customers deploying 4 CPUs per GPU.
- A Georgia Tech and Intel study found CPU-side tool processing accounts for 50–90% of total end-to-end latency in agentic AI workloads.
- Arm CEO Rene Haas forecasts the AI agent era will require 120 million CPU cores per gigawatt, up from 30 million cores per GW for traditional AI data centers.
- Intel CEO Lip-Bu Tan said at Computex 2026 that "for reinforcement learning, orchestration, and agents, the CPU is a much better fit," as the company posted $5.1B in Q1 2026 Data Center and AI revenue (up 22% YoY) and deprioritized consumer chip production to redirect fab capacity to server Xeon parts.
- OpenAI and AWS signed a $38B, 7-year infrastructure partnership in November 2025 covering hundreds of thousands of NVIDIA GPUs with expansion capacity to tens of millions of CPUs for agentic scaling.
- Hugging Face SmolLM2 models (135M–1.7B parameters) enable on-device CPU inference, with the 135M variant fitting entirely in CPU cache on modern smartphones, and domain-specific fine-tuning can achieve performance with 100× less compute.
Why it matters: The CPU-heavy agentic stack reshapes who profits from AI infrastructure: Intel's $5.1B data center quarter and Xeon fab reallocation, plus Arm's projected 4× rise in cores-per-gigawatt, mean buyers must provision orchestration-heavy workloads rather than GPU-dominated training rigs. Local and domain-specific deployments gain a cheaper on-ramp with near-zero marginal hardware cost on already-provisioned server CPUs.
Ask SkimNews



