Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s — SkimNews
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Strata runs the 125-billion-parameter Qwen3.8-Flash-Next model on consumer NVIDIA or AMD GPUs (12 GB+ VRAM) on Windows or Linux, measured at 60 tokens/second for short chats and 32K-token prompts
- The mixture-of-experts architecture activates 10 of 24,576 specialist sub-models per word, with the rest offloaded across GPU, system RAM, and SSD — and a 'guess then check' optimization delivers a stated 1.6-1.8x speedup
- An RTX 3090 (24 GB) is projected to reach 100-140 tokens/second, while a 'Coder' variant retains 91% of the full model's SWE-bench Verified score (weaker on Chinese and other CJK text)
- Strata exposes OpenAI-compatible, Anthropic, and OpenAI Responses API endpoints at http://127.0.0.1:8080, letting Claude Code, Cursor, Codex, and GitHub Copilot route requests to the local model
- The model was compressed by ISTA-DASLab, UkisAI (Swift 1.5), and Unsloth, with Strata built on llama.cpp/ggml and released under the MIT License; the full Q2_0 model is a ~70 GB download that loads 35-55 GB into RAM
Why it matters: Running a 125B-parameter model locally on consumer hardware removes cloud compute costs and data exposure for developers — the Coder variant's 91% SWE-bench Verified retention means coding agents can deliver near-server performance without sending prompts to OpenAI or Anthropic.
Ask SkimNews

