<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk" -->

---
title: NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a...
description: NVIDIA&#x27;s AVO (Agentic Variation Operators) agent architecture achieved a perfect 100.00 RHAE score on the ARC-AGI-3 public benchmark, completing all 183 levels...
canonical: https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents | daily.dev
og:description: NVIDIA&#x27;s AVO (Agentic Variation Operators) agent architecture achieved a perfect 100.00 RHAE score on the ARC-AGI-3 public benchmark, completing all 183 levels...
og:url: https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk
og:image: https://api.daily.dev/og/posts/mwxfl3mwK.png
og:image:alt: NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents

**[NVIDIA Developer](https://daily.dev/sources/nvidiadev)** · 9 min read · 1 upvotes · 0 comments

## Summary

NVIDIA's AVO (Agentic Variation Operators) agent architecture achieved a perfect 100.00 RHAE score on the ARC-AGI-3 public benchmark, completing all 183 levels across 25 environments using 6,624 environment actions with Claude Opus 5, about 12% fewer actions than the comparable VISTA system. AVO was originally developed for autonomous GPU-kernel optimization, where it ran continuously for seven days, explored 500+ optimization directions, and produced attention kernels outperforming cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% on DGX B200 hardware. The same general-purpose agent architecture, featuring persistent memory and a supervisor mechanism, transferred from software/kernel engineering to an unfamiliar interactive reasoning task without domain-specific redesign, suggesting that long-horizon agent capability depends more on system-level machinery than on the underlying model alone.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents>

## Questions this post answers

### What score did NVIDIA's AVO agent achieve on the ARC-AGI-3 public benchmark?

NVIDIA's AVO agent architecture achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels using 6,624 environment actions with Claude Opus 5. For comparison, the VISTA system reportedly used 7,542 environment actions to complete the same 183 levels, meaning AVO used about 12% fewer actions in that cross-system comparison.

_Developers evaluating long-horizon agent architectures can track results like this on daily.dev._

### How much faster did NVIDIA's AVO-optimized attention kernels run compared to cuDNN and FlashAttention-4 on DGX B200?

After a seven-day autonomous run exploring more than 500 optimization directions and producing 40 committed kernel versions, AVO's multihead attention kernels outperformed cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% across evaluated configurations on NVIDIA DGX B200 systems. The agent also adapted the kernel to grouped-query attention in roughly 30 minutes of additional autonomous work.

_Engineers weighing agent-driven kernel optimization against handwritten tuning follow benchmarks like this on daily.dev._

### What are the two key mechanisms that let NVIDIA's AVO agent sustain long-running autonomous tasks?

AVO relies on persistent memory and a supervisor mechanism. Persistent memory carries forward prior implementations, evaluation results, and accumulated reasoning so the agent resumes from its current state instead of reconstructing the search from scratch, while the supervisor monitors the broader trajectory for stagnation and redirects the main agent toward alternative strategies when progress plateaus.

_Teams designing multi-step autonomous agents compare architectural approaches like this on daily.dev._

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#nvidia](https://daily.dev/tags/nvidia), [#gpu](https://daily.dev/tags/gpu), [#claude](https://daily.dev/tags/claude)

[View this post on daily.dev](https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents","url":"https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk"},"datePublished":"2026-08-21T13:02:05.613Z","dateModified":"2026-08-22T21:19:16.876Z","description":"NVIDIA's AVO (Agentic Variation Operators) agent architecture achieved a perfect 100.00 RHAE score on the ARC-AGI-3 public benchmark, completing all 183 levels...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/4346c91c5f6c80470d13951b462cc4ec?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/4346c91c5f6c80470d13951b462cc4ec?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"NVIDIA Developer","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"NVIDIA Developer","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/86e45aab42ba48ce83103d01b1119910","url":"https://daily.dev/sources/nvidiadev"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,nvidia,gpu,claude","timeRequired":"PT9M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"NVIDIA Developer","item":"https://daily.dev/sources/nvidiadev"},{"@type":"ListItem","position":3,"name":"NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-fo-mwxfl3mwk#faq","mainEntity":[{"@type":"Question","name":"What score did NVIDIA's AVO agent achieve on the ARC-AGI-3 public benchmark?","acceptedAnswer":{"@type":"Answer","text":"NVIDIA's AVO agent architecture achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels using 6,624 environment actions with Claude Opus 5. For comparison, the VISTA system reportedly used 7,542 environment actions to complete the same 183 levels, meaning AVO used about 12% fewer actions in that cross-system comparison. Developers evaluating long-horizon agent architectures can track results like this on daily.dev."}},{"@type":"Question","name":"How much faster did NVIDIA's AVO-optimized attention kernels run compared to cuDNN and FlashAttention-4 on DGX B200?","acceptedAnswer":{"@type":"Answer","text":"After a seven-day autonomous run exploring more than 500 optimization directions and producing 40 committed kernel versions, AVO's multihead attention kernels outperformed cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% across evaluated configurations on NVIDIA DGX B200 systems. The agent also adapted the kernel to grouped-query attention in roughly 30 minutes of additional autonomous work. Engineers weighing agent-driven kernel optimization against handwritten tuning follow benchmarks like this on daily.dev."}},{"@type":"Question","name":"What are the two key mechanisms that let NVIDIA's AVO agent sustain long-running autonomous tasks?","acceptedAnswer":{"@type":"Answer","text":"AVO relies on persistent memory and a supervisor mechanism. Persistent memory carries forward prior implementations, evaluation results, and accumulated reasoning so the agent resumes from its current state instead of reconstructing the search from scratch, while the supervisor monitors the broader trajectory for stagnation and redirects the main agent toward alternative strategies when progress plateaus. Teams designing multi-step autonomous agents compare architectural approaches like this on daily.dev."}}]}
```

