DeepSeek V4 Flash costs 99% less than Claude Opus 4.8

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- DeepSeek released V4 Flash at ~$0.28 per million output tokens, roughly 99% cheaper than Anthropic's Claude Opus 4.8 at $25, while matching Opus 4.8 on coding and agentic benchmarks and topping it on Arena.ai's crowdsourced front-end coding leaderboard
- OpenAI cut GPT-5.6 Luna pricing by 80%—from $6 to $1.20 per million tokens—just three weeks after launch, a price SpaceXAI matched when releasing Grok 4.5 for coding, research, and autonomous tasks
- Google released three new Gemini "flash" models focused on efficiency, and Meta quietly reversed its longtime open-weights stance with closed-source Muse Spark 1.1 priced aggressively for developers
- Anthropic remains the clearest premium-pricing holdout, keeping top-tier Claude models at premium pricing and betting developers will pay extra for safety and precision
- Qualcomm VP of AI Vinesh Sukumar told Axios that "intelligent routers"—systems that auto-select the best model per task on capability, speed, and price—could emerge and further weaken any one lab's pricing power
- Former OpenAI head of go-to-market Zack Kass called the trend "diminishing model returns," saying "at some point, the next model doesn't matter to you" as performance gaps between top models shrink
- OpenAI CEO Sam Altman, on the Invest Like the Best podcast, argued volume can offset thinner margins: "We will have so much usage of our models that we do not need to be a gigantically high-margin business to be able to afford model training."
Why it matters: DeepSeek's 28-cent token price versus Anthropic's $25 is a 99% gap—paired with OpenAI's 80% Luna cut, Meta's closed-source pivot, and looming intelligent routers—that threatens frontier labs' ability to recoup the hundreds of billions being poured into AI compute. Whoever owns the routing layer captures pricing power even as individual model margins collapse, reframing the AI race from "best model" to "best dispatcher."




