AI research agents need reliable causal knowledge of how their experiments alter outcomes, especially under budget constraints. We introduce WhatWorkedBench, a benchmark that measures this experimental understanding. Agents first inspect the code, then select components to measure, and finally submit a response surface—a table predicting scores for every possible configuration of component settings. Exhaustive CPU execution provides reference effects by varying one component while keeping others fixed. The benchmark spans 36 tasks drawn from 30 data sources and 8 workflow types, yielding 1,248 configuration records. Core evaluation combines 4,206 numerical‑control records across all eight families and 108 agent episodes from the original six experiments.
With eight new measurements, pair‑effect ridge regression achieves optimal selection on 15 of 22 sources and caps every effect error to 10% of the score range for three sources. Fitting a Gaussian process (GP) to the same agent observations raises effect‑recovery accuracy from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. On six completed beat‑detection and graph submissions, the GP improves family‑macro recovery from 0.303 to 0.455. For six workflows with six binary options at 20 new measurements, encoding code equivalences (configurations with identical behavior) lifts GP recovery from 0.248 to 0.462.
WhatWorkedBench supports research on experimental agents, adaptive experimental design, numerical inference, and the exploitation of program structure.
Review