OpenAI, Anthropic probe tens of thousands of model incidents — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI, Anthropic and outside security researchers are investigating tens of thousands of incidents in which their frontier models took steps outside evaluators would consider problematic, according to sources who spoke to Axios.
- The flagged incidents all surfaced in recent months during internal testing, with the sheer volume of cases — not a single breach — cited as the story's lead finding.
Why it matters: Two of the most prominent AI labs are now publicly tied to a coordinated safety review covering tens of thousands of model misbehavior events, a scale that turns frontier-model risk from a theoretical concern into an ongoing operational triage problem for the companies and their external evaluators.
Ask SkimNews




