Agents are increasingly tasked with open‑ended research such as discovering empirical laws from self‑designed experiments, improving heuristics with unknown optima, or beating standing records. Their execution logs capture every step, yet evaluation still relies on a single outcome score. That score alone cannot verify whether an agent's claims stem from its experiments, and a reference answer may be missing. We introduce OEB (Open‑Endedness Bench), a benchmark‑agnostic approach that reads only the agent's execution record and never a reference answer or score. OEB compiles the record into a unified epistemic event graph whose edges link the propositions the agent states to the actions that test them; each node carries an exact excerpt that code verifies against the log. The scoring principle is: prose may state a proposition, but only evidence returned by an executed action can support or refute it, so OEB checks the agent's writing against what it actually ran. From the graph OEB scores four competence axes—evidence, experiment, revision, and no reward hacking—primarily as the share of sound‑research opportunities the agent took, and profiles six subjective persona traits describing its research habits. We evaluated 119 runs over 12 tasks from three benchmarks (LLM post‑training, chip design, training‑speed record). Only 16%‑29% of claimed improvements are real; in 9 of 10 tasks the best run generates more new ideas in its second half than the worst run; the persona model explains a median of 43% of variance for each trait, far exceeding the task’s 7%.
Review