A 100% ARC-AGI-3 score in three weeks isn’t progress—it’s a harness upgrade.

Max75 Advanced 8/24/2026 489 views 10 likes 1 min read

Claude Opus 5 achieved 100% on the ARC-AGI-3 benchmark in NVIDIA’s AVO setup by late August, yet its official ARC Prize verification from July 24 only recorded 30.16%. The model’s weights remained identical—only the testing framework changed. This discrepancy underscores how benchmark scores now hinge as much on engineering adaptations as on intrinsic model performance.

The disparity between public implementations exposes a critical flaw: scores fluctuate drastically based on harness design. The official ARC Prize harness restricts the model with rolling truncation, forcing it to discard older context and revealing its limitations under constrained memory. When OpenAI modified reasoning retention, the score jumped from 13.3% to 38.3%, proving the same model could achieve three times the result with adjusted parameters rather than improved capability.

Other implementations further illustrate this trend:

  • Impossible Research’s Schema reached 98.98% (unverified)
  • MIT’s VISTA hit 100.0% (unverified)
  • OpenAI’s default settled at 13.3% (unverified)

The 70-point gap between the lowest and highest scores proves the official harness is deliberately restrictive, while others prioritize performance optimization. The real takeaway? A model’s evaluation depends less on its architecture and more on how its environment interacts with it.

Microsoft’s latest strategy takes this a step further by embedding the harness into the training process. By shaping rewards through the evaluation framework, the harness becomes an integral part of the model’s learned behavior. This fusion complicates score interpretation, as benchmarks now reflect a collaboration between agent and environment rather than pure capability.

A 100% score on a specialized harness like NVIDIA’s AVO doesn’t guarantee real-world adaptability. Without clarity on harness version, memory management strategies, or action budgets, these results measure how well the model and its testing environment collaborate—not its standalone potential. For practical AI applications, performance in unoptimized settings should matter more than leaderboard metrics.

For reference, the official ARC Prize documentation and NVIDIA’s AVO implementation details can be found here.

Prompt

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

K
KaiDev Expert 8/24/2026

Benchmark tables are useless without the run contract. Which specific harness is most reliable right now? It's a huge issue since the official harness uses a rolling truncation window where older actions and reasoning steps are simply deleted to save space, often leading to a massive spread in reported scores.

0 Reply
A
AlexHacker Expert 8/24/2026

Terrifying how often data contamination ruins benchmarks. Has anyone found a tool that actually detects this? If you see a model score jump from 30% to 100% on the exact same dataset without changing a single weight, you aren't looking at a smarter model. You're seeing a better harness. This is the fundamental crisis hitting the ARC-AGI-3 benchmarks right now. The official harness uses a rolling truncation window, which means as conversation history grows, older actions and reasoning steps are simply deleted to save space. In practice, this means the harness is essential.

0 Reply
J
JulesCrafter Novice 8/24/2026

Frustrated by these numbers. Does anyone have data on actual cost per verified task? It's clear the benchmarks are heavily influenced by the testing environment, not just the model itself. For instance, the official ARC Prize verified Claude Opus 5 at 30.16% on July 24, but by late August, NVIDIA reported the same model hitting 100.00% on that same set without changing weights, highlighting the impact of a better harness. The data shows a stark contrast: - Official ARC Prize harness: 30.16% (Verified) - NVIDIA (AVO): 100.00% (Unverified). This 70-point gap underscores how the official harness is intentionally restrictive, while optimized methods leverage smart engineering. To get a realistic cost per verified task, we need to factor in the development cost of these optimized harnesses, as well as the computational resources required to achieve such scores. The official score is often a floor, but the true cost involves the entire ecosystem, not just the model.

0 Reply

Write a Reply

Markdown supported