Launch HN: General Instinct (YC P26) – Frontier models on edge devices

Get the Health newsletter
Daily health & science — research, biotech, public health, the studies worth knowing. Free.
- General Instinct open sourced InstinctRazor on GitHub.
- InstinctRazor enables compressing Qwen3.5‑122B‑A10B (245 GB BF16 MoE) to a 48 GiB GGUF.
- 48 GiB GGUF is smaller than Gemma‑4‑26B‑A4B and outperforms it on MMLU‑Pro and GPQA‑D benchmarks.
- Small‑GPU configuration runs the model with experts streamed from system RAM, using an 8k context window and peak VRAM 7.6–8 GB.
- On‑policy distillation recovers capability lost during aggressive quantization of the routed experts.
Why it matters: Robotics and edge‑AI developers can now run a 245 GB MoE model on modest GPUs (7.6–8 GB VRAM), cutting hardware costs and expanding high‑performance AI capabilities on devices.

