Opus 5 blows past Fable 5 and GPT-5.6 Sol on ARC-AGI-3 (The Decoder)
The Decoder (Matthias Bastian), 2026-07-26, reporting ARC Prize‘s analysis of Claude Opus 5 on ARC-AGI-3. The first third-party benchmark read on Opus 5 this spoke has; everything else on that page is Anthropic’s own ratio-only claims.
The numbers (2026-07-26, volatile)
| Benchmark | Opus 5 | Field |
|---|---|---|
| ARC-AGI-3 | 30.2% | GPT-5.6 Sol (Max) 7.8% (previous record); “Fable-class” ~20% per ARC Prize |
| ARC-AGI-2 | 90.4% | matches previous top scores, at slightly higher cost |
| ARC-AGI-1 | 97.5% | matches previous top scores, at slightly higher cost |
| Witness (private) | 43.4 | statistically tied with kimi-k3 and [[claude-fable-5 |
Opus 5 solved five previously unsolved environments, four of them at or above human level; six of the 25 public demo environments are now solved. ARC Prize credits “stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments,” and reports behavior it hadn’t seen before: the model translated tasks into algebraic notation and formulated reflection equations on its own.
The caveat that makes the article worth having
Two independent reasons to discount the 4× headline, both in the piece itself.
The model postdates the benchmark. Opus 5 was developed after ARC-AGI-3 and its format became public, so Anthropic could have targeted the benchmark’s skills through data labeling and RL (labeled reasoning traces, rewarded exploration and self-correction) without training on the exact tasks. Bastian names this as plausible, not established.
A private benchmark shows a much smaller gain. On Witness, Guanghan Ning’s private interactive-puzzle benchmark, Opus 5 scores 43.4 and statistically ties Kimi K3 and Fable 5, improving far less over Opus 4.8 than on ARC-AGI-3. On a puzzle built from common mechanics it stated the hidden rules before its first move; on a game combining rules in a less familiar way it scored below Opus 4.8. Ning reads that as consistent with training on genre-specific data, while noting Witness can’t identify what data Anthropic used, and later clarified that Opus 5 did improve across Witness overall, just far less than on ARC-AGI-3.
Greg Kamradt (ARC-AGI-3) argues the gap doesn’t rule out broader reasoning gains: the conventional puzzle may not test novelty at all, one weak game could be an isolated regression, and Opus 4.8 beat Opus 5 on some ARC-AGI-3 games despite trailing badly overall. Settling it needs detailed results on more unfamiliar tasks.
Ning’s closing analogy is the useful one: coding benchmarks went from a saturated HumanEval to frequently-updated competitions to today’s coding agents, and interactive reasoning will likely follow, with ARC-AGI-3 absorbing the first wave of targeted training before the harder edge cases generalize it.
Tier
T3 — competent tech-press reporting on someone else’s evaluation, with the counter-evidence and both researchers’ caveats included rather than buried. The underlying scores are ARC Prize‘s (public results, replays and benchmarking code) and would be T2 read directly; the Witness figures are one researcher’s private benchmark reported second-hand.
Related
arc-agi · claude-opus-5 · claude-fable-5 · arc-prize · llm-benchmarks · claude-opus-5-announcement · synthesis