ARC Prize Verified

Qwen3.8-27B

Alibaba·Aug 14, 2026·4 reasoning variants

At xhigh effort, Qwen3.8-27B scores 87.5% on ARC-AGI v1 Semi-Private at $0.2152/task and 42.4% on ARC-AGI v2 Semi-Private at $0.4469/task. Medium reasoning performs similarly to low on v1 but scores lower on both v2 sets, despite using more completion tokens. Baseten confirmed that Qwen's chat template adds effort-specific instructions for low and xhigh, but not medium. This may help explain the regression: the settings change the instructions given to the model, rather than simply increasing its thinking budget. Evaluated using dedicated inference on Baseten. Estimated cost uses recorded token usage priced at Alibaba Cloud International's listed rates of $0.425 per million input tokens and $2.55 per million output tokens, using the September 30, 2026 pricing snapshot. This estimate does not represent actual Baseten charges.

ARC-AGI 2 leaderboard

Qwen3.8-27B

Verified scores

VariantARC-AGI-1ARC-AGI-2ARC-AGI-3
XHigh
87.5%
42.4%
—
Medium
68.7%
13.2%
—
Low
69.2%
22.8%
—
None
34.0%
1.5%
—

Tasks & environments

Pass/fail per reasoning level across each benchmark.

ARC-AGI-2 Public Eval

120 tasks
Task
XHigh
Medium
Low
None
✓✗✗✗
✗✗✗✗
✓✗✓✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✓✗
✗✗✗✗
✓✓✓✓
✓✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✓✓✗
✗✗✗✗
✓✓✓✓
✗✗✗✗
✓✗✗✗
✓✓✓✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✓✓✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✓✓✓✗
✗✓✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✓✗✗
✗✗✗✗
✓✗✗✗
✗✗✓✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✓✓✓✗
✓✗✗✗
✗✗✓✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✓✓✓✗
✗✓✓✗
✗✗✗✗
✗✗✗✗
✓✓✓✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✓✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✓✗
✓✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✓✓✓✓
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✓✗✗✗
✓✗✗✗
✓✓✓✗
✗✗✗✗
✓✗✗✗
✗✓✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✓✓✓✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗
✓✗✗✗
✗✗✗✗
✗✗✗✗

ARC-AGI-1 Public Eval

400 tasks
Task
XHigh
Medium
Low
None
✓✓✓✓
✓✓✓✓
✓✓✗✗
✓✓✓✓
✗✓✗✗
✓✗✓✗
✓✓✓✗
✓✓✓✗
✗✗✓✗
✓✗✗✗
✗✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✗✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✗✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✗✗✓
✓✓✓✓
✓✓✓✓
✗✗✗✗
✓✗✗✗
✓✓✓✗
✓✓✓✓
✗✗✗✗
✓✗✗✗
✓✓✓✓
✗✗✗✗
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✗
✗✗✗✗
✗✓✗✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✗✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✗✗✗✗
✓✓✓✗
✓✓✓✗
✗✗✗✗
✗✗✗✗
✗✗✓✗
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✗✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✗✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✓
✗✓✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✗✗✓✗
✓✓✓✗
✓✓✗✓
✓✓✓✓
✓✗✓✓
✓✓✓✓
✓✗✗✗
✓✓✓✓
✓✗✗✗
✓✓✓✓
✓✗✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✗✗✗
✓✓✓✓
✓✓✓✓
✓✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✗✗
✓✓✓✗
✓✗✗✗
✓✓✓✓
✓✓✓✓
✓✗✗✗
✓✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✗✓✗
✓✓✗✗
✗✗✗✗
✓✓✓✗
✓✓✓✗
✓✗✓✗
✓✓✓✓
✓✓✓✗
✓✓✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✗✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✗✗✗
✓✓✓✗
✓✗✗✓
✗✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✗✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✗✗
✓✓✓✗
✗✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✗✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✗✓✓
✓✓✓✗
✓✓✓✗
✗✓✗✗
✗✗✗✗
✓✓✓✗
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✗✓✗
✓✓✓✓
✗✗✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✗✗✗✗
✓✗✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✗✗✗✗
✗✗✗✗
✓✓✓✓
✗✗✗✗
✓✓✗✗
✗✗✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✗✗✗✓
✓✗✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✗✗
✓✓✓✗
✓✓✓✗
✗✗✗✗
✓✗✗✗
✓✓✓✓
✗✓✓✗
✓✗✗✗
✓✓✓✗
✓✗✗✗
✓✓✓✗
✓✓✓✗
✓✓✓✗
✓✓✓✗
✓✓✗✗
✓✓✓✓
✓✓✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✗✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✗✗
✓✓✗✗
✗✗✗✗
✓✓✓✓
✓✗✗✗
✓✓✓✓
✓✓✓✓
✓✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✗✗✗✗
✓✓✓✓
✗✗✗✗
✗✗✓✗
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✗✓✗
✓✓✓✗
✓✓✓✗
✓✓✓✗
✓✓✓✗
✓✗✗✗
✓✓✓✓
✗✓✓✗
✓✓✓✗
✓✗✓✗
✓✓✓✓
✓✗✓✗
✗✗✗✗
✓✗✗✗
✓✓✗✗
✓✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✗✗✗
✓✓✓✗
✓✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✗✗
✓✓✓✗
✓✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✗✗✗
✓✓✓✓
✓✗✗✗
✓✓✓✓
✗✗✓✗
✗✗✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✗✗✗
✓✓✓✗
✓✓✓✗
✓✓✓✗
✗✓✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✗✗
✓✗✗✗
✓✓✗✓
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✗
✗✗✗✗
✓✗✗✗
✓✓✓✓
✗✗✗✗
✓✓✓✗
✗✗✗✗
✓✓✓✓
✓✓✓✗
✗✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✗✓✗
✓✗✗✗
✓✓✓✓
✓✓✓✓
✓✓✓✗
✗✓✓✗
✓✓✓✓
✓✓✓✗
✓✗✓✗
✓✓✗✗
✓✓✓✓
✓✓✓✓
✗✓✓✗
✓✓✓✓
✓✗✓✗
✓✓✓✗
✗✓✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✗✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✓✓✗
✓✓✓✓
✓✗✓✗
✗✗✗✗
✓✓✓✓
✓✓✓✗
✓✓✓✗
✓✓✓✓
✓✗✗✗
✓✓✓✗
✓✓✓✓
✓✓✓✓
✓✗✓✗
✓✗✗✗
✓✗✗✗
✓✓✓✗
✓✓✓✗
✓✓✓✓
✗✗✗✗
✓✓✗✗
✓✗✗✗
✓✗✓✗
✗✓✗✗