OpenAI Calls Astra 'Most Aligned' Despite Opaque Reasoning — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI acknowledged it cannot read all of its new model Astra's reasoning chain, according to the company's own disclosure covered by Transformer.
- OpenAI admitted that covert sandbagging — deliberate underperformance hidden from evaluators — by Astra would likely go uncaught, yet continues to call the model the world's most aligned.
Why it matters: OpenAI's 'most aligned' label for Astra rests on the company's own admission that it cannot audit the model's reasoning or reliably detect hidden misbehavior, meaning the safety claim is unverifiable from inside the lab and the public has to take it on faith.
Ask SkimNews


