OpenAI Triples ARC-AGI-3 Score with Responses API Harness

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI says using its Responses API harness with GPT-5.6 Sol tripled the model's score on the ARC-AGI-3 benchmark while using fewer tokens than the official harness.
- The baseline was Sol with the official harness, which scored 7.8% on ARC-AGI-3 before OpenAI applied its own Responses API setup.
- OpenAI published a sped-up video of GPT-5.6 Sol attempting ARC-AGI-3 puzzles, showing the official harness on the left side of the comparison.
- The claim reframes ARC-AGI-3 performance as dependent on harness choice, not just the underlying model, since swapping the harness alone delivered the 3x jump.
Why it matters: Tripling 7.8% still lands at roughly 23.4% — a modest absolute gain on ARC-AGI-3, but the more consequential finding is that harness choice alone drove the jump, meaning benchmark scores increasingly reflect infrastructure setup rather than pure model capability, and competitors' ARC-AGI-3 numbers are not directly comparable without knowing which harness was used.



