Appendix E · Result registry
Every score, with its protocol attached.
A score alone does not identify an evaluation. Each row keeps it attached to the benchmark edition, system configuration, protocol or split, date, available cost or placement, and supporting evidence.
25
Evaluations
22
Machine results
3
Human baselines
4
With cost data
ARC-AGI-1
13 protocol-labelled evaluations.
Every row here belongs to ARC-AGI-1. Read the benchmark chapter →
| System | Score | Protocol | Date | Cost or rank | Evidence |
|---|---|---|---|---|---|
| Human baselineHuman baseline · arc1-human | 98% | Original ARC human baseline | 2019 | Not recorded | On the Measure of Intelligence ↗ |
| IcecuberResult · arc1-icecuber-2020 | ~20% | 2020 Kaggle solo private result | 2020 | Rank 1 | Kaggle Abstraction and Reasoning Challenge ↗ |
| GreenblattResult · arc1-greenblatt-2024 | ~42% | ARC-AGI-Pub | 2024-06 | Not recorded | ARC Prize 2024 Winners & Technical Report ↗ |
| GreenblattResult · arc1-greenblatt-public-2024 | 50% | Public evaluation | 2024-06 | Not recorded | ARC Prize 2024 Winners & Technical Report ↗ |
| MIT Test-Time TrainingResult · arc1-mit-ttt-2024 | 53% | Public validation, 8B model | 2024-11 | Not recorded | The Surprising Effectiveness of Test-Time Training for ARC ↗ |
| MIT Test-Time Training ensembleResult · arc1-mit-ttt-ensemble-2024 | 61.9% | Ensemble with program synthesis | 2024-11 | Not recorded | The Surprising Effectiveness of Test-Time Training for ARC ↗ |
| The ARChitectsResult · arc1-architects-2024 | 53.5% | 2024 Kaggle private evaluation | 2024 | Not recorded | ARC Prize 2024 Winners & Technical Report ↗ |
| The ARChitectsResult · arc1-architects-public-2024 | 71.6% | Public evaluation | 2024 | Not recorded | ARC Prize 2024 Winners & Technical Report ↗ |
| MindsAIResult · arc1-mindsai-2024 | 55.5% | 2024 season high | 2024 | Not recorded | ARC Prize 2024 Winners & Technical Report ↗ |
| BermanResult · arc1-berman-2024 | 53.6% | ARC-AGI-Pub | 2024-12 | Not recorded | How I got a record 53.6% on ARC-AGI-Pub ↗ |
| OpenAI o3 (low)Result · arc1-o3-low-2024 | 75.7% | 100-task semi-private evaluation, high efficiency | 2024-12 | Not recorded | OpenAI o3 ARC-AGI breakthrough ↗ |
| OpenAI o3 (high)Result · arc1-o3-high-2024 | 87.5% | 100-task semi-private evaluation, approximately 172x compute | 2024-12 | Not recorded | OpenAI o3 ARC-AGI breakthrough ↗ |
| Tiny Recursive ModelResult · arc1-trm-2025 | ~45% | Reported ARC-AGI-1 result | 2025 | Not recorded | ARC Prize 2025 Results & Analysis ↗ |
ARC-AGI-2
7 protocol-labelled evaluations.
Every row here belongs to ARC-AGI-2. Read the benchmark chapter →
| System | Score | Protocol | Date | Cost or rank | Evidence |
|---|---|---|---|---|---|
| Human baselineHuman baseline · arc2-human | ~100% | Every evaluation task solved by at least two humans | 2025 | Not recorded | ARC-AGI-2 ↗ |
| Frontier reasoning modelsResult · arc2-frontier-2025 | ~1–3% | Harness-free frontier-model range | 2025 | Not recorded | ARC-AGI-2 ↗ |
| NVARCResult · arc2-nvarc-2025 | 24.03% | 2025 Kaggle private evaluation; contest constraints | 2025-12-05 | $0.2/task Rank 1 | ARC Prize 2025 Results & Analysis ↗ |
| Tiny Recursive ModelResult · arc2-trm-2025 | ~8% | Reported ARC-AGI-2 result | 2025 | Not recorded | ARC Prize 2025 Results & Analysis ↗ |
| Gemini 3 ProResult · arc2-gemini3-pro-2025 | 31% | ARC Prize verified commercial baseline | 2025-12-05 | $0.81/task | ARC Prize 2025 Results & Analysis ↗ |
| Opus 4.5Result · arc2-opus45-2025 | 37.6% | Thinking, 64k; ARC Prize verified | 2025-12-05 | $2.2/task | ARC Prize 2025 Results & Analysis ↗ |
| Poetiq + Gemini 3 ProResult · arc2-poetiq-gemini3-2025 | 54% | Open refinement harness; ARC Prize verified | 2025-12-05 | $30/task | ARC Prize 2025 Results & Analysis ↗ |
ARC-AGI-3
5 protocol-labelled evaluations.
Every row here belongs to ARC-AGI-3. Read the benchmark chapter →
| System | Score | Protocol | Date | Cost or rank | Evidence |
|---|---|---|---|---|---|
| Human baselineHuman baseline · arc3-human | 100% | First-contact human calibration | 2026 | Not recorded | ARC-AGI-3 ↗ |
| StochasticGooseResult · arc3-stochasticgoose-preview | ~12.58% | 30-day Preview Challenge | 2025-08 | Rank 1 | ARC-AGI-3 Preview: 30-Day Learnings ↗ |
| Blind SquirrelResult · arc3-blind-squirrel-preview | ~6.71% | 30-day Preview Challenge | 2025-08 | Rank 2 | ARC-AGI-3 Preview: 30-Day Learnings ↗ |
| Just-ExploreResult · arc3-just-explore-preview | ~3.64% | 30-day Preview Challenge private leaderboard | 2025-08 | Rank 3 | ARC-AGI-3 Preview: 30-Day Learnings ↗ |
| Frontier LLMsResult · arc3-frontier-harness-free | ~0.1–0.5% | Harness-free frontier | 2025 | Not recorded | ARC-AGI-3 ↗ |
Evaluation receipts
The score is only one field.
A trustworthy result keeps the split, protocol, resource disclosure, artifacts, and claim boundary beside the number. These new ARC-AGI-3 records show the format explicitly.