Four Startups Target the Transformer Bottleneck

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Subquadratic (Miami) claims its sparse attention model SubQ is the first to rival mainstream LLMs on search and coding, selectively processing only the word pairings that matter rather than comparing every token to every other.
- Manifest AI (San Francisco) replaced attention with a "power retention" mechanism that maintains a rolling summary of context, turning open-source coding LLM StarCoder into PowerCoder and releasing Brumby, which it claims rivals versions of Alibaba's Qwen.
- Liquid AI (MIT spinout, Cambridge) builds hybrid models that are 20% transformers and 80% liquid neural networks inspired by worm brains — small enough to run on a $50 Raspberry Pi while matching the performance of LLMs four times their size, including versions of Alibaba's Qwen and Google's Gemma.
- Liquid AI counts Mercedes as a customer, has racked up nearly 34 million downloads, and offers its models free to organizations earning under $10 million in annual revenue.
- Inception (Palo Alto) applies diffusion — the technique behind most image and video generation models — to text, training LLMs to generate whole blocks of text at once rather than word by word.
- The transformer's compute appetite is the cost driver: OpenAI is set to spend $50 billion on computing this year, and the IEA projects total data center electricity consumption will double by 2030.
Why it matters: OpenAI is set to spend $50 billion on computing this year, much of it consumed by the transformer architecture's quadratic-scaling attention. Liquid AI already runs models that match four-times-larger LLMs on a $50 Raspberry Pi for Mercedes — proof that the compute bill driving the AI giants is a design choice, not a physical necessity.
Ask SkimNews




