AI research agents are increasingly employed to search over programs, mathematical constructions, and proofs. Existing systems typically optimize evaluator feedback without proper governance of how that feedback is interpreted, challenged, and reused, which can promote fragile candidates as discoveries and conflate benchmark gains, finite certificates, and theorem‑level claims.
EPOCH introduces an evidence‑governed discovery loop that combines explicit task contracts, typed memory, active falsification, admission checks, and independent replay, ensuring each candidate is evaluated against the strength and scope of the claim it supports.
On the AlgoTune benchmark, EPOCH achieves a mean normalized score of 0.65, substantially exceeding the strongest baseline at 0.53, and attains the highest mean score of 0.57 on the internal Math14 suite. It also shows favorable held‑out behavior under official‑test replay and leads the descriptive aggregate on AgentHPO.
Across ten discovery problems, EPOCH delivers notable task‑specific advances, including executable constructions, optimized algorithms, counterexamples, and proof‑supported results, demonstrating its ability to turn search into concrete progress in mathematical and computational domains.
These findings suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions but also more trustworthy scientific discoveries.
Review