GPT-6 Astra on robot arms — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- GPT-6 Astra succeeded on 19 of 20 "block in bowl" trials using the same YAM arms and Inspect Robots agent policy, against Claude Fable 5.1's 8 of 20 and Fable 5's 1 of 20.
- Fable 5.1 took 6.8 minutes per bowl run at $2.12, versus Astra's 2.5 minutes at $0.94—roughly a third of the time and less than half the cost.
- On the puzzle-piece insertion task, Astra matched Fable 5.1's 2 of 20 completion rate, with both models stalling at the same final insertion step.
- Astra's puzzle runs cost $1.36 versus Fable 5.1's $2.18—cheaper per attempt, but neither model solved the precision-insertion barrier.
- The authors flag limitations: trials ran on different days and weren't interleaved, the bowl comparison used a different rig than the Fable baseline, grading was operator-judged with models identified, and Astra's automatic input caching wasn't factored into its cost estimate.
- The Verge's separate coverage focuses on Sam Altman apologizing for a 'messy' Astra rollout that locked out paying users—an access-controversy angle that leaves the model's actual manipulation capability untouched.
Why it matters: On the bowl task, Astra ran at roughly half Fable 5.1's cost and a third of its time while succeeding nearly 2.4x more often (19/20 vs 8/20). On the precision puzzle, both stalled at the same final step—showing the capability gap is task-specific, not uniform across manipulation.
Ask SkimNews



