Nvidia just showed that the harness, not the AI model, is now the real hero

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Nvidia researchers found that wrapping Claude Opus 5 in a custom harness with good memory management and a "supervisor" agent component pushed it to 100% on the ARC-AGI-3 interactive reasoning benchmark, up from 30% without the harness.
- OpenAI was so frustrated by its models' sub-10% scores on ARC-AGI-3 that it ran its own research, finding two harness tweaks could triple scores — though still far below Nvidia's 100% result.
- Nvidia VP Adel El Hallack argued that for long-horizon agentic tasks, the harness — the tools, runtime, and scaffolding around the model — is a larger determinant of performance than the model itself.
- Microsoft research from April found that all 19 LLMs tested on long-horizon document editing tasks filled documents with errors that would get a human fired.
- Databricks CEO Ali Ghodsi said in July that choosing the wrong harness can 2x AI costs independent of which model is used.
- Nvidia open-sources harness components under its Nemo brand and built a custom harness called Agentic Variation Operators (AVO) to run the benchmark tests.
Why it matters: For enterprise buyers, harness design — not just model selection — can swing agent accuracy from 30% to 100% on long-horizon work and can 2x costs, per Databricks' CEO. Nvidia's push to open-source scaffolding via Nemo positions it to monetize the stack around agents, while El Hallack explicitly tied the argument to OpenAI slowing training over security breaches.
Ask SkimNews




