A The ARC Atlas

Appendix I · From claims to traces

Observe how learning unfolds.

Scores are endpoints. Runs reveal the path: which action produced information, when a model became predictive, where a prior failed, what was remembered, and whether the evidence transfers beyond the public set.

342
Human replays indexed
25
Public environments
145
Completed runs
458
Study participants

Run Observatory

Twenty-five worlds. One uneven learning landscape.

The official public corpus contains 342 first-run human replays. This index keeps the complete environment-level accounting visible and links every row to the playable task. Search or sort without losing the no-JavaScript table underneath.

25 environments

EnvironmentPlaysSolvesCompletionTop scoreOpen
ar25105 50% 8 levelsPlay ↗
bp35142 14% 9 levelsPlay ↗
cd82118 73% 6 levelsPlay ↗
cn04126 50% 6 levelsPlay ↗
dc22114 36% 6 levelsPlay ↗
ft09104 40% 6 levelsPlay ↗
g50t143 21% 7 levelsPlay ↗
ka59103 30% 7 levelsPlay ↗
lf52114 36% 10 levelsPlay ↗
lp855415 28% 8 levelsPlay ↗
ls20136 46% 7 levelsPlay ↗
m0r0117 64% 6 levelsPlay ↗
r11l1010 100% 6 levelsPlay ↗
re86115 45% 8 levelsPlay ↗
s5i5114 36% 8 levelsPlay ↗
sb26125 42% 8 levelsPlay ↗
sc251510 67% 6 levelsPlay ↗
sk48147 50% 8 levelsPlay ↗
sp80122 17% 6 levelsPlay ↗
su15133 23% 9 levelsPlay ↗
tn36146 43% 7 levelsPlay ↗
tr87126 50% 6 levelsPlay ↗
tu93139 69% 9 levelsPlay ↗
vc33106 60% 7 levelsPlay ↗
wa30145 36% 9 levelsPlay ↗

The Atlas indexes the complete official aggregate. Full step-by-step frames remain in the official dataset; this site embeds only traces whose artifacts are openly attributable. Official dataset and methodology ↗

Human Learning Atlas

Difficulty has a shape.

Aggregate success ranges from 100% on r11l to 14% on bp35. That spread is the first diagnostic: a benchmark average can hide worlds with very different onboarding cliffs and learning burdens.

Teaching traces

Three traces, three different lessons.

These are evidence-backed reading guides, not reconstructed frame data. They show what to look for before you open the official step-by-step corpus.

r11l

A low-friction world

10 of 10 participants solved it

  1. Probe an available action
  2. Recognize the repeated response
  3. Reuse the mechanic across six levels

High solvability does not mean zero learning. It means the environment yields a useful rule quickly and consistently.

Inspect the official evidence ↗
cd82

The onboarding cliff

8 of 11 solved; two stalled before level 3

  1. Explore early controls
  2. Encounter the level-2 bottleneck
  3. Transfer the discovered mechanic through later levels

Per-level progression distinguishes a hard first abstraction from uniformly hard execution.

Inspect the official evidence ↗
re86

Success is not efficiency

5 finished; only 4 reached 100% efficiency

  1. Acquire the mechanic
  2. Complete every level
  3. Compare action count with the median-human baseline

A binary solve hides whether the agent learned economically. RHAE preserves that distinction.

Inspect the official evidence ↗

Failure Museum

A failure is useful when it changes the next experiment.

The diagnosis map separates an observed symptom from its proposed mechanism. Each repair is deliberately framed as a test, not a guaranteed cure.

ObservedExplains a local visual effect but cannot predict the next state

DiagnoseNarrative recognition without a causal transition model

Test nextRecord the transition, make a falsifiable prediction, then probe the smallest disagreement

ObservedImports a familiar game rule despite contradictory observations

DiagnoseA prior is treated as fact rather than a hypothesis

Test nextList competing hypotheses and choose an action whose outcomes separate them

ObservedSolves one level, then rediscovers the same mechanic later

DiagnoseUseful knowledge is not consolidated into durable memory

Test nextStore a compact rule, its evidence, and the conditions under which it applies

ObservedEventually completes the environment after many redundant actions

DiagnosePlanning is disconnected from the human-normalized action budget

Test nextTrack information gained per action and stop experiments once the decision changes

Ablation Board

Put the components on trial.

The 2026 ablation study matters because it refuses a comforting story. Executable persistence alone is not uniformly helpful; simplification usually helps; exact verification is consistently strongest in the main comparisons but costs more.

TreatmentPersistent modelSimplificationExact replayFinding
Textual baselineNoNoNoCan beat a flexible executable model in some settings
Flexible executable modelYesNoNoPersistence alone is not reliably beneficial
+ scheduled simplificationYesYesNoImproves 3 of 4 main model-effort settings
+ fixed interface and verificationYesYesYesRanks first in all 4 main settings; uses more resources

Reproducibility Index

Can another team inspect the claim?

Grades summarize artifact completeness, not scientific importance. “Open code” is only one column; prompts, run artifacts, environment definition, and held-out evaluation determine whether a result can be audited.

SystemCodePromptsRun artifactsEnvironmentHeld-out evidenceGrade
EWMOpenOpenOpenDocumentedNot testedA−
OPINE-WorldPaper-linkedDescribedPartialDocumentedNot testedB
EWM ablationPaper-linkedTreatments specifiedReportedDocumentedNot testedB+
Milestone #1 agentsOpen submissionsVariesVariesCompetitionSemi-privateB
Human corpusDatasetProtocol open342 replays25 publicFirst-run humansA

The next useful result is not merely a higher number. It is a higher number with a clearer causal story and a stronger receipt.

→ Turn the open questions into experiments

YouTube first · local synths follow

ARC Radio

01 / 12 🦉 8-Bit Chiptune Playlist 🦉 Retro Video Game Music for Nostalgic Vibes YouTube · external stream

The 4 requested YouTube selections play first and require a network connection; their titles refresh from YouTube when they load. 8 original AI-composed retro-game loops follow and are generated live in your browser. Audio keeps playing when you close this panel and stops only when you press Pause.

Field notes · reader review

Help improve this guide

Found a wrong score, broken link, missing paper, or unclear passage? Tell us what you noticed.

How useful is it? optional
- / 5
What kind of note? optional

No account, no tracking. Sent straight to the maintainer.