OpenAI's Ultrafast Runs GPT-5.6 Sol at 14x Speed

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI launched Ultrafast, a new mode that runs GPT-5.6 Sol at 14x standard processing speed and delivers up to 750 output tokens per second.
- Cerebras powers the Ultrafast preview under its chip partnership with OpenAI, with access currently limited to a small group of customers as capacity grows.
- OpenAI is positioning Ultrafast for enterprise workflows including incident response, customer service and support, financial market analysis, and e-commerce.
- Anthropic offers a Claude fast mode, but the article notes it does not deliver the kind of speed OpenAI is offering with Ultrafast.
- OpenAI framed Ultrafast as a departure from the trade-off of "choosing a smaller or more specialized model," pointing instead toward "more useful work per second."
Why it matters: By hitting 750 tokens/second on a frontier model, OpenAI is directly targeting latency-sensitive enterprise workflows — incident response and financial market analysis — where smaller competitors' models previously dominated. The Cerebras partnership gives OpenAI a concrete speed differentiator against Anthropic's slower fast mode, though the preview's limited customer base means the real measure is whether capacity scales.
Ask SkimNews




