AVO, Nvidia's general purpose coding agent, hit 100% on ARC AGI 3, a public interactive reasoning benchmark from the ARC Prize team.
Nvidia said its general-purpose coding agent, AVO, completed all 183 levels across all 25 public environments of ARC-AGI-3, a 183-level interactive reasoning benchmark designed to test general intelligence rather than memorized skill. The result, posted to Nvidia's AI account on X, is the first 100% score Nvidia has reported on the test.
According to Nvidia, AVO "figured out what to do with no instructions, explicit rules, or stated goals." Nvidia frames the work in a company blog post and a mechanism paper on arXiv describing "Agentic Variation Operators for Autonomous Evolutionary Search." AVO used 12% fewer actions than Anthropic's Opus 5 running without the AVO scaffold.
The 100% is Nvidia's only score. The ARC Prize methodology and technical report define a clean run as one where harness and external scaffolding are explicitly excluded. Nvidia describes AVO as a general-purpose coding agent system with persistent memory, supervisor monitoring, and tool use — a scaffold-based architecture that operates within the benchmark's environment but is itself an external scaffold relative to the model.
The "no instructions, explicit rules, or stated goals" framing describes the benchmark's baseline test condition — the starting state the benchmark is specifically designed to create — rather than a condition AVO uniquely achieved.
A Hacker News discussion flagged that end-to-end completion time is not disclosed, and no Codex-with-a-goal baseline comparison has been published. The benchmark's creator, the ARC Prize team, has not commented on the score.