OpenAI's Ultrafast Mode Runs GPT-5.6 Sol at 14x Speed

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI rolled out "Ultrafast," a new mode for its GPT-5.6 Sol model that runs at 14x standard processing speed and delivers up to 750 output tokens per second, according to a Thursday blog post.
- OpenAI framed the launch as a departure from trading capability for speed, writing: "Until now, getting real-time speed typically meant choosing a smaller or more specialized model."
- Cerebras is powering the Ultrafast preview through its chip partnership with OpenAI, with access currently limited to a small group of customers.
- OpenAI is positioning Ultrafast for corporate use cases including incident response, customer service and support, financial market analysis, and e-commerce.
- Anthropic has launched a competing "fast mode" for Claude, though OpenAI explicitly contrasted that offering's speed against its own 14x throughput claim.
- OpenAI said Ultrafast is in preview and that availability will expand "as capacity grows."
Why it matters: By leaning on Cerebras silicon rather than scaling GPT-5.6 Sol itself, OpenAI is making speed — not just capability — a buying criterion for enterprise customers in incident response, financial analysis, and support, while publicly drawing a contrast with Anthropic's Claude fast mode that the company says doesn't match the 750-tokens-per-second pace.
Ask SkimNews




