DeepSeek V4 Flash API Hits Public Beta at $0.14/MToken

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- DeepSeek rolled out the V4 Flash API in public beta, with Bloomberg citing the company's claim of benchmark scores "far surpassing" its own V4 Pro Preview
- DeepSeek-V4-Flash-0731 is published on Hugging Face, where its model card states it outperforms the larger DeepSeek-V4-Pro
- DeepSeek-V4-Flash-High scored 1586 on the Frontend Code Arena at $0.14 input / $0.28 output per million tokens — "best performance-per-dollar of any model in its class"
- DeepSeek kept V4-Flash's architecture and size identical to the preview version; the V4-Pro API, App, and Web models remain unchanged, with V4-Pro's official release "coming ASAP"
- DeepSeek answered OpenAI's overnight price cut with this release, according to Nikkei Asia, intensifying the AI price war
- A hacker has already used DeepSeek AI to autonomously attack vulnerable servers (BleepingComputer) — a concrete security dimension sitting alongside the pricing story in the aggregator's coverage
Why it matters: DeepSeek is pricing V4-Flash-High at $0.14 input / $0.28 output per million tokens while claiming the best performance-per-dollar in its class. The timing — arriving immediately after OpenAI's price cut, per Nikkei — shows Chinese labs are now setting the floor on commodity inference pricing. Enterprise buyers gain immediate negotiating leverage against any vendor still charging multiples of that rate.




