Microsoft's MDASH Scores 95.95% on CyberGym at Half Cost

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Microsoft launched MAI-Cyber-1-Flash, its first cybersecurity-specific AI model, inside MDASH, reporting a 95.95% CyberGym score when paired with GPT-5.4 — but the headline figure belongs to the MDASH system, not the model on its own.
- The new configuration replaces 80% of MDASH's existing models and is claimed to cost 50% less than Microsoft's current best mix (GPT-5.4, GPT-5.4 mini, GPT-5.4 Codex); however, Microsoft did not disclose token use, call volume, latency, or compute allocation behind the comparison.
- CyberGym's public leaderboard did not list the 95.95% result as of July 28, 2026, still showing Microsoft's earlier 88.4% May 12 MDASH submission, and Microsoft's June 96.55% figure used different scoring criteria so the three numbers are not directly comparable.
- MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137 billion total parameters, 5 billion active, and a 256,000-token context window, fine-tuned from MAI-Code-1-Flash; it is designed to handle up to 90% of MDASH tasks with GPT-5.4 reserved for the hardest 10%.
- On ExploitGym, which tests whether agents can turn supplied vulnerabilities into working code-execution exploits, the model scored zero across the kernel, userspace, and browser categories.
- Access is limited to approved MDASH customers through an Azure AI Foundry private preview, and this is the first announced scenario for Project Perception, Microsoft's broader agentic security coordination system scheduled for public preview on August 3.
- Taesoo Kim, Microsoft's VP of agentic security, framed the launch with: "The model is one input, the system around it is the product."
Why it matters: Microsoft itself frames MDASH — not MAI-Cyber-1-Flash — as the product, so the 95.95% number is a system-level claim that has not been independently reproduced or listed on CyberGym's public leaderboard. Enterprise customers evaluating the 50% cost-savings pitch must do so without disclosed token use, call volume, or compute allocation, and the model's zero score on ExploitGym exploit-generation tasks sits in tension with the headline benchmark narrative.


