GPT-6 Astra on robot arms — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- GPT-6 Astra placed a red block into a bowl in 19 of 20 trials on YAM robotic arms, versus Claude Fable 5.1's 8 of 20 and Fable 5's 1 of 20, at 2.5 minutes and ~$0.94 per run
- Claude Fable 5.1 took 6.8 minutes and ~$2.12 per bowl trial, more than double Astra's time and cost on the same task under the same Inspect Robots agent policy
- On the round blue puzzle-piece insertion, Astra and Fable 5.1 each completed just 2 of 20 trials, stalling at the same final step at $1.36 vs $2.18 per run
- The benchmark's authors flagged limitations: trials run two days apart rather than interleaved, a different rig for the bowl comparison, operator-known grading, and list-price costs that don't discount Astra's automatic input caching
- Other outlets angled the rollout as controversial: The Verge reported Sam Altman apologizing for a messy GPT-6 Astra rollout that locked out paying users, while TechMeme/Fortune highlighted OpenAI revising evaluation metrics in ways that appear to favor Astra
Why it matters: For agent-policy buyers choosing between frontier models for robot control, GPT-6 Astra is roughly 2.4× more reliable and less than half the per-run cost than Fable 5.1 on the bowl task—but the identical stall on the puzzle insertion exposes a shared, fine-grained failure mode neither vendor's marketing addresses. With OpenAI separately under fire for revised metrics and a rollout that locked out paying users, a third-party benchmark that pins down a real capability gap carries more editorial weight than vendor-supplied numbers.
Ask SkimNews


.png)
