ARC Prize Verified

GPT-6 Luna

OpenAI·Sep 22, 2026·12 harness configurations

GPT-6 Luna's best Semi-Private scores are 86.7% on v1 and 59.3% on v2, both at max reasoning. On v3, it reaches 0.2% with the Standard harness at medium reasoning and 0.6% with the Provider Adapter harness at max reasoning.

ARC-AGI 3 leaderboard

GPT-6 LunaGPT-6 Luna - Provider Adapter

Verified scores

VariantARC-AGI-1ARC-AGI-2ARC-AGI-3(Standard harness)ARC-AGI-3(Provider Adapter harness)
Max
86.7%
59.3%
0.10%
0.59%
XHigh
73.0%
41.9%
0.16%
0.51%
High
70.3%
31.4%
0.18%
0.38%
Medium
61.0%
18.1%
0.19%
0.32%
Low
37.7%
4.6%
0.03%
0.24%
None
8.8%
0.0%
0.03%
0.05%

Tasks & environments

Pass/fail per reasoning level across each benchmark.

ARC-AGI-3 Public Demo - Standard harness

25 environments

ARC-AGI-3 Public Demo - Provider Adapter harness

25 environments

ARC-AGI-2 Public Eval

120 tasks
Task
Max
XHigh
High
Medium
Low
None

ARC-AGI-1 Public Eval

400 tasks
Task
Max
XHigh
High
Medium
Low
None