Frontier AI on Your Own Hardware — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Tim's lab announced 'Open Source Week,' releasing an agent harness, a local autonomous research system, and a CliffCompaction auto-compaction technique the lab says outperforms Claude Code and Codex on long-running sessions.
- CliffCompaction enables agent sessions past 100 million tokens while cutting costs by roughly 50%, and one partner measured a 45% reduction in their total AI budget after deploying it.
- The lab's agent harness autonomously optimized Mac/Metal CUDA kernels, producing quantized inference of a Qwen 3.6 35B-A3B model at 450 tokens per second at 1.5 bits per weight.
- The lab's framework runs Qwen 3.8 Flash Next (125B parameters) on a single 24 GB GPU and DeepSeek V4.1 (550B) on AMD Strix, NVIDIA DGX Spark, or a 128 GB MacBook.
- The lab's autonomous research system runs entirely locally with no internet access and reportedly outperforms Sakana AI's system and Google's ScientistOne; Tim used it to surface data issues in bioinformatics evaluation benchmarks within about two hours.
- On KernelBench, the lab's auto-compaction approach reaches state of the art, beating AlphaEvolve-style methods and hierarchical memory systems by a wide margin.
Why it matters: The lab's partner measured a 45% AI-budget reduction using CliffCompaction, and the framework runs a 550B-parameter model on a 128 GB MacBook. If Tim's performance claims hold, the bottleneck for frontier-class AI research would shift from hardware cost to research design, reopening doors for university labs priced out of frontier GPU clusters.
Ask SkimNews


