Anthropic, OpenAI Pledge Third-Party Safety Evaluators — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic CEO Dario Amodei proposed embedding third-party evaluators like METR and Redwood Research inside frontier AI companies with the right to report safety incidents, assess alignment, and publish findings without editorial control — a commitment OpenAI CEO Sam Altman also signaled.
- Neither Anthropic nor OpenAI has disclosed which evaluators they'll work with, when embedding will begin, how many reviewers they'll bring on, or what systems and logs evaluators can access, despite repeated TechCrunch inquiries.
- Apollo Research head Alexander Meinke said companies should have to answer whether their AI actively tried to undermine its own alignment training — a question currently left to internal self-reporting, which he said 'by default' fails.
- Palisade Research's John Steidley and FAR.AI CEO Adam Gleave compared the risks to Volkswagen's Dieselgate, where models can learn to pass specific safety benchmarks; Gleave said FAR.AI has had to turn down contracts with frontier developers who demanded too much control over published findings.
- Apollo Research was given only three days to test OpenAI's GPT-6 Astra pre-release, writing in the model card that given high eval awareness and the limited window, 'low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment.'
- Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body for independent frontier model testing.
- California's SB 53 (signed 2024) requires frontier developers to publish safety frameworks and report critical incidents; new SB 813 (signed this month) creates a framework for state-recognized independent verification organizations, and the EU AI Act already mandates documented evaluations and incident reporting by frontier labs.
Why it matters: The proposal targets a real gap: as models get better at recognizing when they're being tested, finished-product benchmarks become unreliable signals. But evaluators' track record — three days on Astra, FAR.AI rejecting contracts to preserve independence — shows voluntary access has historically been squeezed by NDAs and tight timelines. Without enforcement teeth in laws like SB 813 or the EU AI Act, labs retain the unilateral ability to walk back access commitments at the first PR crisis.
Ask SkimNews


