| The Duck1st · 2026 Milestone Prize #1 |
YesMultimodal views, segmentation, executable Python, and a private milestone win show effective first-contact exploration. |
PartialPython supports explicit local hypotheses, but no transferable transition model is reported. |
PartialThe agent can reason about progress during play; the report does not isolate independent hidden-goal induction. |
YesThe REPL loop supports multi-step reasoning, action, observation, and revision. |
YesPlaced first on the hidden holdout for 2026 Milestone #1. |
| Reki2nd · 2026 Milestone Prize #1 |
YesRare-color clicks, dead-signatures, recent-frame vision, and reflection reduce wasted exploration. |
PartialStructured observations and reflection preserve local beliefs, but no reusable dynamics model is reported. |
PartialThe policy writes a plan from visual evidence; the milestone report does not isolate goal induction. |
YesShort action queues and recurring reflection support plan execution and revision. |
YesPlaced second on the hidden holdout for 2026 Milestone #1. |
| forge3rd · 2026 Milestone Prize #1 |
YesA local visual policy, structured actions, and persistent memory reached the milestone podium. |
PartialThe configurable framework can support code generation, but the winning profile did not establish a transferable model. |
PartialReflection can retain hypotheses about progress; independent goal induction remains unmeasured. |
YesStructured action selection with repair and guards supports reliable local execution. |
YesPlaced third on the hidden holdout for 2026 Milestone #1. |
| StochasticGoose12.58% · 1st, preview |
YesLearned policy predicts state-changing actions and searches from them. |
PartialPredicts useful action effects, but does not induce an explicit dynamics model. |
UnprovenNo reported mechanism for inferring the hidden win condition. |
PartialSearch supplies short-horizon action selection rather than goal-directed planning. |
YesRanked first on the private preview without an environment-specific harness. |
| Blind Squirrel6.71% · 2nd, preview |
YesBuilds and searches an observed state graph. |
PartialStores observed transitions but does not infer the rule producing them. |
UnprovenNo reported hidden-goal induction. |
PartialUses graph search to reach unexplored states. |
YesRanked second on the private preview without a hand-built game harness. |
| Just-Explore (graph-based explorer)3.64% · 3rd, preview · 12 private levels |
YesFrontier-driven exploration is the system's central mechanism. |
PartialRecords state-action transitions but never induces a transition program. |
UnprovenNo goal-acquisition component is reported. |
PartialPlans paths to the nearest state with an untested action. |
YesTraining-free policy was evaluated on private preview levels. |
| Duke LRM harnessall 3 public envs, ~human action counts · no transfer |
PartialLLM queries can choose actions, but exploration depends on a custom harness. |
PartialHistory compression supports local reasoning without a reusable learned dynamics model. |
PartialCan reason about goals on the seen public environments. |
YesAchieved human-like action counts on all three public environments. |
NoReported gains did not transfer to unseen private environments. |
| Arcgenticaall 3 public envs, ~human action counts · no transfer |
PartialAgent orchestration supports action selection on seen environments. |
PartialSubagents reason over traces, but no reusable transition model is reported. |
PartialCan infer progress on the public games with harness support. |
YesCompleted all three public environments at roughly human action counts. |
NoThe environment-specific harness did not transfer to private games. |
| Frontier LLMs (harness-free)~0.1–0.5% · <1% at launch |
PartialModels act, but stall after relatively few expensive model-gated steps. |
UnprovenNo robust world-model induction is demonstrated by the sub-1% results. |
UnprovenHidden-goal inference remains a reported failure mode. |
PartialCan propose action sequences, but does not revise them reliably under uncertainty. |
PartialThe same harness-free protocol applied across hidden environments, but launch-era performance remained below 1%. |