A The ARC Atlas

Appendix G · Visual index · 31 plates

The book's visual evidence, in one place.

Browse every chart, map, leaderboard, and data table in source-page order. Each plate matches its chapter appearance and links back to the surrounding argument.

31
Figures & tables
8
Source pages
15
Figures
16
Tables

The complete visual set

Read the field by its shapes.

Figures appear in source-page orderFigures and tables use separate counters, so a plate number identifies its kind, not its position in this gallery.. Use each source link to return to the argument and annotations around the visual.

TABLE 01

Three generations of ARC

How task format, action, feedback, scoring, and failure mode change from ARC-AGI-1 to ARC-AGI-3.

Atlas · The Thesis ↗

Question answered: What changed when ARC moved from static transformations to instruction-free interactive worlds?

These are different evaluation regimes. Score percentages should not be compared as one continuous series.

Task format, interaction, feedback, scoring, and exposed failure across ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3.
Dimension ARC-AGI-12019–2024 ARC-AGI-22025– ARC-AGI-32025–
Task Infer one transformation from a few examples. Infer harder, more compositional transformations. Discover and solve an unseen interactive game.
Observation Static input/output grids. Static input/output grids with fewer shortcuts. A changing 64×64, 16-color world.
Agent action Submit up to two output grids. Submit up to two output grids. Press a key, undo, or click a cell each turn.
Feedback Demonstration pairs only. Demonstration pairs only. The next frame after every action; no instructions.
Scoring Exact-match pass@2. Exact-match pass@2 under a cost cap. Completion plus Relative Human Action Efficiency.
Failure exposed Few-shot rule induction. Composition, refinement, and efficiency. Exploration, world modeling, goal acquisition, and planning.

Sources: ARC-AGI-1 · ARC-AGI-2 · ARC-AGI-3.

FIG 01

The ARC-AGI timeline

Every catalogued work, placed by year and primary theme.

Atlas · Timeline ↗
ARC-AGI paper timeline The catalogued works arranged by year and primary theme. Measure & method The o3 break The redesign Frontier Benchmark design Agentic frontier Program synthesis Test-time training Neural / compression Milestone results Foundations Critique & survey 2019 2020 2021 2022 2023 2024 2025 2026
FIG 02

The web of ideas

How the papers build on, answer, and compete with one another.

Atlas · The Web ↗
ARC-AGI idea graph Connections among the catalogued ARC-AGI works. EFFICIENT · GENERALIZATION ·
FIG 03

The human-vs-AI gap

Best reported AI score against the human baseline for each benchmark.

Atlas · The Gap ↗
ARC-AGI-12019
87.5%
98%
ARC-AGI-22025
24.03%
100%
ARC-AGI-32026
12.58%
100%
Best AI Human baseline ▨ the gap that remains
FIG 04

The climb on ARC-AGI-1

The record’s rise from Icecuber to o3.

Atlas · The Gap ↗
ARC-AGI-1 frontier climb Selected record scores over time compared with the human baseline. 0 25 50 75 100 human ≈ 98% 20% Icecuber 2020 42% Greenblatt Jun 2024 53% MIT TTT Nov 2024 53.6% Berman Dec 2024 75.7% o3 (low) Dec 2024 87.5% o3 (high) Dec 2024
FIG 05

Cumulative works in the Atlas

How the selected catalog grows over time.

Atlas · By the Numbers ↗
Cumulative works The cumulative number of catalogued works over time. 0 7 14 211920212223242526 24 worksthe 2024 rush
FIG 06

Influence × year

Every work by year, influence, and primary theme.

Atlas · By the Numbers ↗
Influence versus year Every catalogued work arranged by year and editorial influence rating. 1 2 3 4 51920212223242526 Influence
FIG 07

Works per year

Selected works by publication year.

Atlas · By the Numbers ↗
1
19
2
20
21
22
2
23
8
24
6
25
5
26
TABLE 02

Works by theme

The thematic composition of the Atlas.

Atlas · By the Numbers ↗
Synthesis
10
Test-time
7
Benchmark
5
Agentic
5
Neural
5
Foundations
5
Milestone
4
Critique
3
TABLE 03

Works by format

Papers, reports, posts, code, and other source formats.

Atlas · By the Numbers ↗
Papers
18
Reports
2
Blog posts
3
Code
1
TABLE 04

Recurring authors

Names appearing on two or more works.

Atlas · By the Numbers ↗
Chollet
3
Ellis
2
Kamradt
2
Knoop
2
Landers
2
Pu
2
Rodionov
2
TABLE 05

Most-connected works

Works with the most cross-references to other Works.

Atlas · By the Numbers ↗
TABLE 06

The chronological catalog

All works ordered by year with theme and influence.

Atlas · The Chronicle ↗
#YearWorkThemeInfluence
01№01 2019 On the Measure of IntelligenceFrançois Chollet · paper Benchmark 5 02№07 2020 ARC 2020 Kaggle 1st-Place Solution (Icecuber)Johan S. Wind · code Milestone 5 03№18 2020 DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learningEllis et al. · paper Foundations 4 04№14 2023 Hypothesis Search: Inductive Reasoning with Language ModelsWang et al. · paper Synthesis 4 05№19 2023 Large Language Models as General Pattern MachinesMirchandani et al. · paper Foundations 3 06№06 2024 OpenAI o3 Breakthrough High Score on ARC-AGI-PubARC Prize · blog Milestone 5 07№04 2024 ARC Prize 2024: Technical ReportChollet et al. · report Benchmark 4 08№08 2024 The Surprising Effectiveness of Test-Time Training for Abstract ReasoningAkyürek et al. · paper Test-time 4 09№10 2024 Getting 50% (SoTA) on ARC-AGI with GPT-4oGreenblatt et al. · blog Synthesis 4 10№11 2024 How I got a record 53.6% on ARC-AGI-Pub using Sonnet 3.5 with Evolutionary Test-time ComputeJeremy Berman · blog Synthesis 4 11№12 2024 Combining Induction and Transduction for Abstract ReasoningLi et al. · paper Synthesis 4 12№13 2024 Searching Latent Program Spaces (LPN)Bonnet et al. · paper Neural 3 13№15 2024 CodeIt: Self-Improving Language Models with Prioritized Hindsight ReplayButt et al. · paper Synthesis 3 14№02 2025 ARC-AGI-2: A New Challenge for Frontier AI Reasoning SystemsChollet et al. · paper Benchmark 5 15№09 2025 Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of PerspectiveFranzen et al. · paper Test-time 5 16№17 2025 Graph-Based Exploration for ARC-AGI-3 Interactive Reasoning TasksRudakov et al. · paper Agentic 4 17№24 2025 Don't throw the baby out with the bathwater: How and why deep learning for ARCCole et al. · paper Critique 4 18№16 2025 ARC-AGI Without Pretraining (CompressARC)Liao et al. · paper Neural 3 19№23 2025 Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGIPfister et al. · paper Critique 2 20№03 2026 ARC-AGI-3: A New Challenge for Frontier Agentic IntelligenceChollet et al. · paper Agentic 5 21№22 2026 Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?Sergey Rodionov · paper Agentic 5 22№05 2026 ARC Prize 2025: Technical ReportChollet et al. · report Benchmark 4 23№20 2026 Executable World Models for ARC-AGI-3 in the Era of Coding AgentsSergey Rodionov · paper Agentic 4 24№21 2026 OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3Courtis et al. · paper Agentic 4
FIG 08

The price of an extra action

How ARC-AGI-3’s efficiency term falls as an agent spends more actions than a human.

ARC-AGI-3 · RHAE ↗

Question answered: How sharply does ARC-AGI-3 punish actions beyond the human baseline?

This curve includes the official 1.15 per-level cap. The per-game completion cap, level weighting, and benchmark aggregation are separate parts of the final score.

Relative Human Action Efficiency by agent action count The score is capped at 115 percent for faster-than-human action counts, is 100 percent at the human action count, 25 percent at twice the human count, and 4 percent at five times the human count. 0% 25% 50% 75% 100% 115% 0.75× agent actions ÷ human actions →

human actions → 25% efficiency

level efficiency = min(1.15, (human actions / agent actions)²)

Anchor values for the curve
Agent actions0.75×
Efficiency term115%100%25%11.1%6.25%4%

Source: ARC-AGI-3 scoring methodology. Retrieved 2026-07-28.

TABLE 07

ARC-AGI-3 preview leaderboard

Selected preview scores, with the harness-free frontier shown for comparison.

ARC-AGI-3 · Solutions ↗
1
StochasticGoose · Tufa Labs
12.58%
2
Blind Squirrel
6.71%
3
Frontier LLMs · harness-free
<1%
TABLE 08

Capabilities supported by published evidence

Published evidence for exploration, modeling, goal inference, planning, and harness-free generalization.

ARC-AGI-3 · Solution Space ↗

Question answered: Which parts of the ARC-AGI-3 loop does each reported solver actually demonstrate?

Statuses summarize published evidence, not an intrinsic capability score. “Generalize” is an added cross-cutting diagnostic, not one of the benchmark’s four core capabilities.

Evidence coverage across current ARC-AGI-3 solution attempts.
SystemExploreModelInfer goalPlanGeneralize
The Duck1st · 2026 Milestone Prize #1 YesMultimodal views, segmentation, executable Python, and a private milestone win show effective first-contact exploration. PartialPython supports explicit local hypotheses, but no transferable transition model is reported. PartialThe agent can reason about progress during play; the report does not isolate independent hidden-goal induction. YesThe REPL loop supports multi-step reasoning, action, observation, and revision. YesPlaced first on the hidden holdout for 2026 Milestone #1.
Reki2nd · 2026 Milestone Prize #1 YesRare-color clicks, dead-signatures, recent-frame vision, and reflection reduce wasted exploration. PartialStructured observations and reflection preserve local beliefs, but no reusable dynamics model is reported. PartialThe policy writes a plan from visual evidence; the milestone report does not isolate goal induction. YesShort action queues and recurring reflection support plan execution and revision. YesPlaced second on the hidden holdout for 2026 Milestone #1.
forge3rd · 2026 Milestone Prize #1 YesA local visual policy, structured actions, and persistent memory reached the milestone podium. PartialThe configurable framework can support code generation, but the winning profile did not establish a transferable model. PartialReflection can retain hypotheses about progress; independent goal induction remains unmeasured. YesStructured action selection with repair and guards supports reliable local execution. YesPlaced third on the hidden holdout for 2026 Milestone #1.
StochasticGoose12.58% · 1st, preview YesLearned policy predicts state-changing actions and searches from them. PartialPredicts useful action effects, but does not induce an explicit dynamics model. UnprovenNo reported mechanism for inferring the hidden win condition. PartialSearch supplies short-horizon action selection rather than goal-directed planning. YesRanked first on the private preview without an environment-specific harness.
Blind Squirrel6.71% · 2nd, preview YesBuilds and searches an observed state graph. PartialStores observed transitions but does not infer the rule producing them. UnprovenNo reported hidden-goal induction. PartialUses graph search to reach unexplored states. YesRanked second on the private preview without a hand-built game harness.
Just-Explore (graph-based explorer)3.64% · 3rd, preview · 12 private levels YesFrontier-driven exploration is the system's central mechanism. PartialRecords state-action transitions but never induces a transition program. UnprovenNo goal-acquisition component is reported. PartialPlans paths to the nearest state with an untested action. YesTraining-free policy was evaluated on private preview levels.
Duke LRM harnessall 3 public envs, ~human action counts · no transfer PartialLLM queries can choose actions, but exploration depends on a custom harness. PartialHistory compression supports local reasoning without a reusable learned dynamics model. PartialCan reason about goals on the seen public environments. YesAchieved human-like action counts on all three public environments. NoReported gains did not transfer to unseen private environments.
Arcgenticaall 3 public envs, ~human action counts · no transfer PartialAgent orchestration supports action selection on seen environments. PartialSubagents reason over traces, but no reusable transition model is reported. PartialCan infer progress on the public games with harness support. YesCompleted all three public environments at roughly human action counts. NoThe environment-specific harness did not transfer to private games.
Frontier LLMs (harness-free)~0.1–0.5% · <1% at launch PartialModels act, but stall after relatively few expensive model-gated steps. UnprovenNo robust world-model induction is demonstrated by the sub-1% results. UnprovenHidden-goal inference remains a reported failure mode. PartialCan propose action sequences, but does not revise them reliably under uncertainty. PartialThe same harness-free protocol applied across hidden environments, but launch-era performance remained below 1%.
DemonstratedPartial evidenceNot yet demonstratedNot demonstrated

Sources: ARC-AGI-3 Technical Report · ARC-AGI-3 Preview: 30-Day Learnings · ARC Prize 2026: ARC-AGI-3 Milestone Prize #1.

FIG 09

The climb and the resets

The frontier score across all three ARC-AGI benchmarks.

Methods & Techniques · The Climb ↗
The climb and resets Selected frontier scores across the three ARC-AGI benchmark generations. 0 25 50 75 100'20'21'22'23'24'25'26 human ≈ 98% 20% 42% 55.5% 87.5%ARC-AGI-1 3% 24.03%ARC-AGI-2 0.5% 12.58%ARC-AGI-3
ARC-AGI-1ARC-AGI-2ARC-AGI-3 Human ≈ 98–100%
TABLE 09

ARC-AGI-1 selected results

Reported scores across public, private, and ARC-AGI-Pub splits.

Methods & Techniques · ARC-1 ↗
1
OpenAI o3 · OpenAI
87.5%
2
MindsAI · Cole, Osman, Hodel
55.5%
3
Berman — Evolutionary · Jeremy Berman
53.6%
4
The ARChitects · Franzen, Disselhoff, Hartmann
53.5%
5
MIT — Test-Time Training · Akyürek et al. · MIT
53%
6
Greenblatt (GPT-4o) · Ryan Greenblatt · Redwood
42%
7
Icecuber · Johan S. Wind, solo
20%
TABLE 10

ARC-AGI-2 selected results

Reported scores and per-task costs for the 2025 season's leading systems.

ARC-AGI-2 · Solutions ↗
1
Gemini-3-Pro harness · application-layer refinement
54%
2
2025 Kaggle winner · top private entry
24.03%
3
Tiny Recursive Model (TRM) · weight-space refinement
8%
4
Frontier reasoning models · o3 · o1-pro · Claude 3.7
3%
FIG 10

Selected ARC-AGI-2 score–cost frontiers

Frontiers among selected systems, computed separately for each evaluation class.

Methods & Techniques · ARC-2 ↗

Question answered: Which selected ARC-AGI-2 systems buy more score for their reported inference cost?

Frontier status is computed only within the same declared evaluation class and only among the systems shown. Contest, commercial-model, and refinement results share axes for orientation but are not treated as one controlled evaluation.

Selected ARC-AGI-2 scores versus reported cost per task Cost uses a logarithmic horizontal axis. Higher scores and lower costs are better. Frontier status is computed separately for each evaluation class among the selected systems shown. 0% 15% 30% 45% 60% $0.1 $1 $10 $100 NVARC 24.03% · $0.2 Gemini 3 Pro 31% · $0.81 Opus 4.5 37.6% · $2.2 Poetiq + Gemini 3 Pro 54% · $30 reported cost per task · log scale →
2025 contest · private evalARC Prize verified · commercial modelsARC Prize verified · refinement harness Non-dominated within shown evaluation class
Underlying score, cost, protocol, and frontier status.
SystemScoreCost/taskProtocolStatus
NVARC24.03%$0.22025 Kaggle private evaluation; contest constraintsFrontier among shown peers
Gemini 3 Pro31%$0.81ARC Prize verified commercial baselineFrontier among shown peers
Opus 4.537.6%$2.2Thinking, 64k; ARC Prize verifiedFrontier among shown peers
Poetiq + Gemini 3 Pro54%$30Open refinement harness; ARC Prize verifiedFrontier among shown peers

Snapshot 2025-12-05. Source: ARC Prize 2025 Results & Analysis.

TABLE 11

ARC-AGI-3 selected results

Preview and harness-free systems shown with their evaluation context.

Methods & Techniques · ARC-3 ↗
1
StochasticGoose · Tufa Labs
12.58%
2
Blind Squirrel
6.71%
3
Just-Explore · Rudakov, Shock & Cowley
3.64%
4
Frontier LLMs (harness-free) · major labs
<1%
FIG 11

Method genealogy: recurring mechanisms

A reading map of combinations and methodological analogies across ARC solver families.

Methods & Techniques · Genealogy ↗

Question answered: Which mechanisms recur or combine across ARC solver families?

A curated reading map of recurring mechanisms. Solid edges mark a combination; dashed edges mark a close methodological analogy. Neither establishes citation volume or causal descent.

Reading map of recurring ARC solver mechanisms Four equal-width lanes compare explicit program synthesis, neural adaptation, iterative refinement, and interactive world search from 2020 to 2025. Solid lines show combinations and dashed lines show methodological analogies. Explicit programs Neural adaptation Refine + verify Search a world DSL search2020 · Synthesis DreamCoder2020 · Synthesis LLM patterning2023 · LLM Hypothesis search2023 · Synthesis Program sampling2024 · Synthesis CodeIt2024 · Synthesis TTT2024 · Test-time Product of Experts2024 · Test-time LPN2024 · Test-time Evolutionary synthesis2024 · Evolution BARC ensemble2024 · Synthesis o3 guided search2024 · Search CompressARC2025 · Compress Graph exploration2025 · Search
CombinationMethodological analogy
  1. DSL searchHypothesis searchmoves the explicit rule into language + code
  2. Hypothesis searchProgram samplingscales execution-filtered sampling
  3. Program samplingEvolutionary synthesisadds iterative selection and repair
  4. DSL searchBARC ensemblesupplies an explicit-program branch
  5. Program samplingBARC ensemblesupplies neural program induction
  6. LLM patterningTTTcontrasts frozen prediction with per-task adaptation
  7. TTTProduct of Expertscombines adaptation with multi-view agreement
  8. TTTLPNmoves adaptation into a latent program
  9. DreamCoderCodeItshares a learn-from-search loop
  10. CodeItEvolutionary synthesisturns candidate feedback into refinement
  11. Program samplingo3 guided searchreplaces broad sampling with a learned guide
  12. o3 guided searchGraph explorationcompares search over traces with search over world states
FIG 12

The technique map

Methods placed by induction vs. transduction and frozen vs. test-time adaptation.

Methods & Techniques · The Map ↗
Technique map ARC-AGI techniques arranged by induction versus transduction and frozen versus test-time adaptation. ← Predict the answer (transduction) Search for a rule (induction) → Adapt at test time ↑ ↓ Frozen model ~20% ~42% 56.75% 30% ~15% concept 53% 53.5% ~78%* 53.6% ~20% 87.5% 3rd zero-shot
SynthesisTest-timeEvolutionCompressSearchLLM
TABLE 12

The generalization boundary

What each evaluation surface establishes—and what remains outside its evidence.

ARC-AGI-3 · Evidence boundary ↗
Evidence surfaceWhat it provesWhat it does not prove
Observed transitionsThe model can replay evidence already seenIt predicts unseen states
Unseen states in one gameWithin-environment dynamics generalizeThe representation transfers
Unseen public gamesSome cross-game adaptationResistance to public-set tuning or contamination
Semi-private / private gamesHeld-out benchmark transferBroad agentic intelligence outside ARC
FIG 13

Agent architecture atlas

Five recurring ways ARC-AGI-3 agents represent, explore, remember, verify, and plan.

Winning Solutions · Architecture atlas ↗
01

Directed state graph

Blind Squirrel / Just Explore

Represent
Observed states and transitions
Explore
Search unvisited graph edges
Remember
Explicit state graph
Verify
State equality and replay
Plan
Graph search

Evidence boundary: Limited

02

Learned action effects

StochasticGoose

Represent
CNN transition/value estimates
Explore
Prioritize likely progress
Remember
Replay buffer
Verify
Observed reward and transition
Plan
Informed search

Evidence boundary: Model-dependent

03

VLM tool harness

The Duck / Reki / forge

Represent
Images, grids, segments, notes
Explore
Model-selected probes
Remember
Summaries, queues, eviction
Verify
Guards and visual checks
Plan
Short-horizon tool use

Evidence boundary: Public/private milestone evidence

04

Executable world model

EWM

Represent
Persistent Python simulator
Explore
Resolve model contradictions
Remember
Code plus transition log
Verify
Exact replay
Plan
Simulate before acting

Evidence boundary: Public set only

05

Ontology-guided model

OPINE-World

Represent
Object-centric program
Explore
Prioritize ontology error
Remember
Program plus counterexamples
Verify
CEGIS replay
Plan
Model-based

Evidence boundary: Public set only

TABLE 13

Research experiment board

Questions converted into interventions, measurements, falsifiers, and required receipts.

Open Problems · Experiment board ↗
QuestionInterventionMeasureFalsifierRequired receipt
Do executable models improve held-out generalization?Text-only vs executable vs verified executable agentsPrivate-set RHAE, solves, actions, tokens, wall timeNo consistent private-set gain after cost matchingFresh workspaces, prompts, code, transition logs, per-game seeds
Does ontology error choose better experiments?Random, novelty, disagreement, and ontology-error action selectionActions to first correct predictive modelNo gain across unseen mechanicsPer-action hypotheses, predictions, observations, posterior changes
Which memories transfer without leaking game identity?No memory, episodic traces, compressed mechanics, learned retrievalCross-environment efficiency under identity maskingTransfer disappears on private familiesMemory snapshots, retrieval logs, contamination audit
Where do humans spend their learning actions?Align first-run human and agent progression by levelActions to insight, post-insight execution slope, abandonment pointAgent and human curves differ only by a constant scaleAnonymized step replays and a declared event-labeling protocol
TABLE 14

Evaluation receipts

Scores shown with split, protocol, cost, artifacts, and claim boundary.

Evaluations · Evidence receipts ↗
ClaimResultSplitProtocolCostArtifactsBoundary
Human public corpus145 solves / 342 plays25 Public DemoFirst-run, one attempt90-minute sessionsStep replaysHuman cohort; not AI performance
EWM · GPT-5.5 high58.12% mean per-game RHAE25 public gamesFresh agent and workspace per gameNot normalized hereCode and run artifactsPrivate validation untested
OPINE-World78.4 action efficiency25 public gamesNo per-game trainingNot normalized herePaper-linked implementation evidencePublic set only
Verified EWM · gpt-5.6-sol~99% RHAE25 public gamesExploratory follow-upHigher-resource verificationAblation reportPublic-set saturation; model postdates games
FIG 14

Human learning across public environments

Completion rates for all 25 public environments, preserving the unequal shape of first-run difficulty.

Run Observatory · Human learning ↗
FIG 15

Failure diagnosis map

Observable behavior connected to a likely mechanism and a discriminating repair.

Run Observatory · Failure museum ↗

ObservedExplains a local visual effect but cannot predict the next state

DiagnoseNarrative recognition without a causal transition model

Test nextRecord the transition, make a falsifiable prediction, then probe the smallest disagreement

ObservedImports a familiar game rule despite contradictory observations

DiagnoseA prior is treated as fact rather than a hypothesis

Test nextList competing hypotheses and choose an action whose outcomes separate them

ObservedSolves one level, then rediscovers the same mechanic later

DiagnoseUseful knowledge is not consolidated into durable memory

Test nextStore a compact rule, its evidence, and the conditions under which it applies

ObservedEventually completes the environment after many redundant actions

DiagnosePlanning is disconnected from the human-normalized action budget

Test nextTrack information gained per action and stop experiments once the decision changes

TABLE 15

Executable-world-model ablation board

Four nested treatments separate persistence, simplification, and exact replay verification.

Run Observatory · Ablation board ↗
TreatmentPersistent modelSimplificationExact replayFinding
Textual baselineNoNoNoCan beat a flexible executable model in some settings
Flexible executable modelYesNoNoPersistence alone is not reliably beneficial
+ scheduled simplificationYesYesNoImproves 3 of 4 main model-effort settings
+ fixed interface and verificationYesYesYesRanks first in all 4 main settings; uses more resources
TABLE 16

Reproducibility index

A compact inventory of code, prompts, artifacts, environment definition, and held-out evidence.

Run Observatory · Reproducibility ↗
SystemCodePromptsRun artifactsEnvironmentHeld-out evidenceGrade
EWMOpenOpenOpenDocumentedNot testedA−
OPINE-WorldPaper-linkedDescribedPartialDocumentedNot testedB
EWM ablationPaper-linkedTreatments specifiedReportedDocumentedNot testedB+
Milestone #1 agentsOpen submissionsVariesVariesCompetitionSemi-privateB
Human corpusDatasetProtocol open342 replays25 publicFirst-run humansA

YouTube first · local synths follow

ARC Radio

01 / 12 🦉 8-Bit Chiptune Playlist 🦉 Retro Video Game Music for Nostalgic Vibes YouTube · external stream

The 4 requested YouTube selections play first and require a network connection; their titles refresh from YouTube when they load. 8 original AI-composed retro-game loops follow and are generated live in your browser. Audio keeps playing when you close this panel and stops only when you press Pause.

Field notes · reader review

Help improve this guide

Found a wrong score, broken link, missing paper, or unclear passage? Tell us what you noticed.

How useful is it? optional
- / 5
What kind of note? optional

No account, no tracking. Sent straight to the maintainer.