DevOps / SRE / Platform · 21.08.2026, 13:16 UTC
Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 21.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
Nvidia first introduced its Agentic Variation Operators (AVO) general-purpose coding agent system in late March 2026. The company has now unravelled the architecture and the system-level mechanisms that enable AVO to sustain long-running autonomous work and applied it to the open platform advanced AI reasoning benchmark ARC-AGI-3.
In a team blog released on Friday, a five-person team of Nvidia software engineers, machine learning specialists and AI research interns describe how AVO elevated Claude Opus 5 from a reported 30.2% model baseline score on ARC-AGI-3, to 100% when run as part of the complete AVO agent system.
“[This result] shows that system design – not model capability alone – can unlock frontier-level long-horizon performance,” wrote the team.
What tasks does Nvidia’s AVO handle?
Evolving from the core DNA of what constitutes an agent harness, AVO handles the agentic architecture tasks associated with inspecting and editing code, running commands, consulting documentation, and validating its work through execution. AVO’s distinguishing focus is said to be “sustained in-context autonomous operation” across extended, multistep tasks on long horizons.
Non-profit AI research and benchmarking body ARC Prize used an analysis post this July to report the 30.2% score for Claude Opus 5 at high reasoning effort on the public set of environments and tasks on the ARC-AGI-3 system. According to ARC Prize, this “demonstrates strong logical reasoning” and, when Claude Opus 5 was run at Max reasoning effort, it scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2 …