Kog is going deeper to squeeze more inference out of GPUs

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Kog, a French startup, hit Hacker News in May with a demo showing 3,000 tokens per second on a 2-billion-parameter model called Laneformer 2B using AMD MI300X and NVIDIA H200 GPUs — standard datacenter hardware rather than purpose-built chips like Cerebras's
- CEO Gaël Delalleau said the demo generated 200 tangible business leads, with software engineering as the expected first use case, since veteran Claude Code users "sometimes have to wait hours" for results and Anthropic charges a premium for Fast Mode
- Kog claims "30x faster LLM inference" through deep-level GPU optimization but hasn't yet proven it on large language models; Delalleau targets a 10x-speed demo on a major model in September to unlock Series A funding
- Kog's team of 11 spends weeks to months optimizing each new GPU, applying a methodology Delalleau says is shaped by his background as a four-time DEF CON CTF finalist and École Polytechnique-trained solid-state physicist
- Kog is backed by Bpifrance and French Tech 2030, with Varsity VC — led by former Stribe cofounder Kamel Zeroual — co-leading the seed round, and is also supported by Scaleway
- Kog competes with French peer ZML, which bypasses Nvidia's CUDA for hardware-agnostic inference, but Delalleau says his approach more closely resembles Stanford's Hazy Research in its low-level GPU focus
Why it matters: Inference speed is now a billable line item (Anthropic charges a premium for Fast Mode) and a workflow bottleneck (Claude Code users wait hours). If Kog delivers meaningful speedups on GPUs enterprises already own, it undercuts the value proposition of purpose-built inference chips like Cerebras — but the company still has to prove 10x speed on real LLMs at a September milestone before its Series A.
Ask SkimNews



