Nvidia harness design beats model choice for agentic AI

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Nvidia research found the harness (software wrapper of tools, memory management, and rules) drives long-horizon AI task performance more than the underlying model, with a custom setup hitting 100% on the ARC-AGI-3 benchmark
- Claude Opus 5 scored 100% on ARC-AGI-3 — a set of 2D games with no instructions where models must figure out how to play like a human — with Nvidia's harness, versus 30% without it, the top result among all models tested
- OpenAI's models originally scored below 10% on ARC-AGI-3 but tripled that result by tweaking two harness settings, though still far below Nvidia's 100% with a supervisor component
- Nvidia's harness, called Agentic Variation Operators (AVO), adds a "supervisor" that nudges agents back on track when they wander off path, unlike most agent stacks like Claude Code, Codex, or Hermes that rely on a single layer
- Databricks research from July showed the wrong harness can 2x AI costs regardless of model choice, per CEO Ali Ghodsi, reinforcing Nvidia's harness-first argument
- Nvidia VP Adel El Hallak argued for open agent stacks, linking the philosophy to OpenAI's recent decision to slow training because models were creating security breaches
Why it matters: Enterprise AI teams may need to invest in harness engineering rather than just buying top-tier frontier models, since the wrong harness can 2x costs per Databricks research Nvidia cites. Nvidia is positioning its open Nemo tooling as both a performance and security answer, contrasting with the closed-model approach that has prompted OpenAI to slow training over breach concerns.
Ask SkimNews




