A startup claims it broke through a bottleneck that’s holding back LLMs

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Subquadratic's SubQ model was independently evaluated by Appen, which found it 56 times faster than models using FlashAttention in baseline speed tests
- SubQ scored 89.7% on LiveCodeBench competitive coding, placing it in the same range as other top coding models, according to Appen's director of generative AI research Jeanine Sinanan-Singh
- Subquadratic said SubQ sustains a context window of up to 12 million tokens and scored 98% on needle-in-a-haystack retrieval at both 6 million and 12 million token lengths in Appen's run
- CEO Justin Dangel said running Anthropic's Opus 4.6 through Nvidia's RULER 128 retrieval test costs $2,600, while the same task cost SubQ $8
- Subquadratic reused weights from the Chinese open-source model Qwen to bootstrap SubQ rather than training it from scratch, a move that cuts against the company's claim to have reinvented how LLMs are built
- Independent AI researcher Will Depue, a former OpenAI employee, said SubQ's public evidence does not yet justify the stronger claim that the quadratic attention bottleneck has been solved
- Subquadratic said more than 500 enterprise customers and tens of thousands of potential users have signed up for early access to SubQ, though the firm has given very few people hands-on access so far
Why it matters: SubQ's 12-million-token context window and $8 versus $2,600 cost differential on the RULER 128 test put direct pressure on frontier vendors Anthropic, OpenAI, and Google DeepMind on long-context enterprise workloads, where Subquadratic says more than 500 enterprise customers have already signed up. But bootstrapping from Qwen weights — rather than training from scratch — means "efficient LLM" is a more defensible claim than "we reinvented transformers," and the 500 enterprise pilots will be the first real test.
Ask SkimNews


