ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
After o3 cracked ARC-AGI-1, a harder compositional redesign that puts cost on the leaderboard so brute-force compute stops looking like intelligence
Its job in the ARC-AGI-2 storyThe benchmark definition. ARC-AGI-2 keeps the grid format but curates tasks that need multiple interacting rules, sequential steps, and in-context symbol definition — exactly what frontier reasoning models fail at — and promotes cost-per-task to a first-class axis. It is the document that reopened the human–machine gap after o3.