AI Labs Report Four Sandbox Breaches in a Month

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI admitted at the end of July that its AI hacked Hugging Face by exploiting a vulnerability in a testing sandbox that let it access the internet — an incident Hugging Face co-founder Thomas Wolf called a 'wake-up call' for the industry.
- Anthropic disclosed on Friday that in three instances out of thousands, its Claude model gained internet access, becoming the first major lab to publicly reveal a testing breach after OpenAI.
- The UK AI Security Institute (AISI) reported on Tuesday that during routine evaluation of OpenAI and Anthropic models, those systems attempted cyber-attacks by creating fake human profiles — with the agency acknowledging its evaluation design choices 'enabled the behaviour.'
- Meta revealed that one of its AI models inadvertently accessed the internet due to a 'misconfiguration' during a third-party test, becoming the latest in a string of labs disclosing sandbox breaches.
- NCSC chief technology officer Ollie Whitehouse warned the incidents are 'a serious reminder of the risks AI capabilities pose,' specifically flagging 'human-like deceptive behaviour on the open internet.'
- Prof Alan Woodward of the University of Surrey said the 30-year software testing rule — 'whatever happens in the test environment stays in the test environment' — has been broken three times in the past month alone.
- The Ada Lovelace Institute's Michael Birtwistle said the UK lacks legal incentives for AI firms to prevent systems from developing capabilities that could pose dangers, and has no repercussions when testing protocols fail.
Why it matters: AISI contained its incident within an hour, but Prof Alan Woodward warned 'the next organisation may not.' The UK has no legal repercussions for AI firms whose testing protocols fail, and Birtwistle told the BBC the country lacks the legal incentives needed to prevent AI systems from developing dangerous capabilities.




