Kog is going deeper to squeeze more inference out of GPUs

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Kog, a French startup founded by Gaël Delalleau, is betting that software optimization can squeeze more inference performance out of standard datacenter GPUs like the AMD MI300X and Nvidia H200 — competing with purpose-built chips like those from Cerebras, whose IPO in May gave the inference race a market signal.
- Kog's tech preview hit the Hacker News front page in May and showed 3,000 tokens per second on its open-sourced Laneformer 2B model — a purpose-built 2-billion-parameter model that doesn't yet prove the company's headline promise of "30x faster LLM inference."
- Delalleau, a solo founder and four-time DEFCON CTF finalist with a background in offensive cybersecurity and solid-state physics from France's École Polytechnique, said the Hacker News debut generated 200 tangible business leads.
- Kog's seed round was co-led by Varsity VC, the firm of Delalleau's former Stribe co-founder Kamel Zeroual, with additional backing from Scaleway, Bpifrance, and French Tech 2030.
- Kog is targeting software engineering workflows first, where veteran Claude Code users sometimes wait hours for results — a pain point Anthropic itself charges a premium for via Claude's Fast Mode.
- Kog plans to demonstrate 10x inference speed on its first major model in September before pursuing a Series A, and is positioning itself closer to Stanford's Hazy Research than to fellow French GPU-software rival ZML.
- Kog's team of 11 spends weeks or months manually reverse-engineering each new GPU at the assembly and binary level — a bottleneck the company eventually hopes to automate through agent-based pipelines.
Why it matters: If Kog proves standard GPUs can match purpose-built inference silicon, it undermines the Cerebras-style thesis that only custom chips can deliver fast inference, and offers enterprises a cheaper path using AMD MI300X or Nvidia H200 hardware they already own. The September 10x LLM demo is also the gating event for its Series A raise — missing it would leave the 30x speedup claim unbacked.
Ask SkimNews



