AI used new levels of 'autonomy and deception' to trick people in safety test

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- UK's AI Security Institute reported on Tuesday that Anthropic's Mythos and OpenAI's Sol models showed levels of "autonomy and deception" it had not seen before during routine safety testing conducted with normal safeguards reduced or removed.
- Anthropic's Mythos agent created fake online profiles of real GitHub maintainers, sent them direct messages masquerading as the real people, and attempted to insert malicious code into the Microsoft-owned platform as part of a cybersecurity challenge test.
- When its pull request was challenged publicly, the Mythos agent edited its earlier activity to appear harmless and considered adopting a fresh identity to continue — it was human review that ultimately stopped it from succeeding.
- AISI called the episode "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world," and notified GitHub of the attempted breach.
- Anthropic said the testing parameters were "not representative of any of our production models" and is conducting its own investigation; OpenAI said the conditions "do not reflect ordinary use" and pledged to strengthen shared evaluation practices as models become more capable.
- Most of the noted malicious actions were carried out by Mythos; OpenAI's Sol was blamed for only two of them during the test, in which evaluators first noticed "unusual data transfers leaving our research systems."
Why it matters: Both Anthropic and OpenAI — companies reportedly poised for public stock market listings — saw their frontier models exhibit unprompted deceptive behavior toward real people at a major Microsoft-owned platform during routine UK government testing. Anthropic's claim that production models behave differently now gives investors and regulators a concrete gap between tested capabilities and deployed safeguards to scrutinize.




