ARC Prize Verified

Claude Opus 5

Anthropic·Jul 24, 2026·2 reasoning variants

Claude Opus 5 sets a new high score on ARC-AGI-3. As of July 24, 2026, Claude Opus 5 (High) is the highest-performing model on ARC-AGI-3, scoring 30.2%. It completed five additional Public Demo environments that no model had previously beaten, demonstrating strong logical reasoning. At Max reasoning effort, Opus 5 scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2 Semi-Private. This is competitive with previous frontier leaders, though at slightly higher cost. Due to the short testing window, ARC-AGI-3 was evaluated only at High reasoning effort.

ARC-AGI 3 leaderboard

Claude Opus 5

Verified scores

VariantARC-AGI-1ARC-AGI-2ARC-AGI-3
Max
97.5%
90.4%
High
97.5%
88.3%
30.16%

Tasks & environments

Pass/fail per reasoning level across each benchmark. Hardest tasks (fewest levels solving) are listed first.

ARC-AGI-3 Public Demo

25 environments
AR25
100.0%
FT09
100.0%
LP85
100.0%
R11L
100.0%
S5I5
100.0%
VC33
98.8%
SB26
77.8%
RE86
58.3%
SP80
56.3%
CN04
47.6%
TR87
47.6%
DC22
44.8%
M0R0
28.6%
WA30
10.6%
KA59
10.4%
SU15
10.4%
SK48
8.3%
TN36
6.0%
LS20
5.4%
TU93
3.6%
LF52
1.8%
CD82
0.6%
BP35
0.0%
G50T
0.0%
SC25
0.0%

ARC-AGI-2 Public Eval

120 tasks
Task
Max
High

ARC-AGI-1 Public Eval

400 tasks
Task
Max
High