Open-Weight AI Cyber Gap Narrows to 4-7 Months — SkimNews
.png)
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- AISI found open-weight AI models now lag frontier closed models' cyber capabilities by 4-7 months, down from 6-10 months through most of 2025, in its first public open/closed weight gap analysis
- GLM-5.2 matches Opus 4.6 and GPT-5.3-Codex (released 4 months before it) on narrow cyber tasks, and performs comparably to Opus 4.5 on AISI's "The Last Ones" cyber range
- DeepSeek V4-Pro matches Opus 4.5 on narrow tasks but falls below Sonnet 4.5 on "The Last Ones"; AISI says its occasional task refusals were "easily circumvented by a small number of repeat attempts"
- AISI said open-weight model evaluations were "largely unimpeded by safeguards," warning of a narrow window before frontier cyber capabilities reach widely accessible open models
- OpenAI's GPT-5.6 Sol is the state of the art in cyber per Greg Brockman, outperforming Anthropic's Mythos 5 on AISI's evaluation; Ramez Naam noted "everyone now has frontier hacking capabilities"
- Moonshot's Kimi K3 took #1 on the Frontend Code Arena and scored 88.3 on Terminal Bench 2.1 (per VentureBeat); AISI intends to test it once weights are released
Why it matters: The 4-7 month gap means defenders have less runway before frontier hacking capabilities reach less-safeguarded open models — AISI explicitly noted the evaluations were "largely unimpeded by safeguards" and DeepSeek's occasional refusals were "easily circumvented." The Financial Times framed the finding as Chinese AI models narrowing the cyber gap with US rivals, sharpening the geopolitical stakes for export-control debates already underway.
Ask SkimNews
.png)
.png)

