ARC‑AGI‑3 evaluates agents in interactive environments where rules and goals must be inferred from observation. We introduce Kepler, an open‑source harness that treats hypotheses as executable world models and validates them via retrospective transition checks and conditional prediction checks.
Using a frozen Claude Opus 5 configuration, Kepler achieved a server‑verified 100.00 RHAE on all 25 public games, without per‑game model selection or score‑conditioned reruns.
On 181 of 183 completed levels the final Opus attempt used no more actions than the median‑human baseline. The retained board runs comprised 8,256 environment actions, of which 7,292 occurred in scored levels.
Local provider‑session records amount to 858.0 million tokens, with a 97.37% cache‑read rate, costing roughly $777.72 at the September 1 2026 API list‑equivalent rates.
We also report three evaluation failures: source‑code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner.
A single‑game observation case study showed that animation frames contained task‑relevant information absent from settled text grids.
Across the final Claude Opus 5 and GPT‑5.6 Sol boards, 48 of 50 game‑model cells reached 100.
These findings indicate that public‑set scores alone have limited discriminative value and motivate first‑attempt, cost‑conditioned, and verification‑aware reporting.
Review