Startups Chase What Comes After the Transformer

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Subquadratic (Miami) claims its SubQ model is the first sparse attention mechanism that rivals mainstream LLMs on search and coding tasks, selectively calculating which token pairings matter rather than comparing every word against every other; the company says thousands have signed up for its waitlist.
- Manifest AI (San Francisco) built a mechanism called "power retention" that keeps a rolling summary of the context window instead of tracking everything in it, adapted the open-source StarCoder into PowerCoder, and released Brumby, which it claims rivals parts of Alibaba's Qwen.
- Liquid AI (MIT spinout, Cambridge, MA) pairs transformers with liquid neural networks inspired by worm brains, producing hybrid LFMs that are 20% transformer and 80% liquid; the firm says its models match competitors four times their size, can run on a $50 Raspberry Pi, serve Mercedes, and have racked up nearly 34 million downloads.
- Inception (Palo Alto) applies diffusion — the technique behind most image- and video-generation models — to text, producing whole blocks of words at once rather than one token at a time, which co-founder Stefano Ermon says lets a big transformer "predict many tokens at the same time."
- The transformer itself, introduced in Google's 2017 "Attention Is All You Need" paper, is now the bottleneck: dense attention can demand 50 million multiplications for a 10,000-word document, and OpenAI is set to spend $50 billion on computing this year while the IEA projects data-center electricity use will double by 2030.
- Liquid AI's model-design process is itself run by another AI that searches combinations of liquid, convolutional, and transformer networks — an approach CEO Ramin Hasani frames as a deliberate bid to escape the "20 watts of power" gap between human brains and modern LLMs.
Why it matters: The four startups are each targeting the same bottleneck: the quadratic cost of attention, which the article links to OpenAI's $50 billion in annual compute spending and the IEA's projection of doubled data-center electricity by 2030. If even one of these architectures — sparse attention, power retention, hybrid liquid networks, or diffusion — displaces the dense transformer at scale, the cost and energy curve of every major LLM shifts, and the compute moat that protects today's giants thins.
Ask SkimNews




