AI research agents combine prior knowledge, public sources and experimental feedback to generate useful results. The Discovery Certification Protocol (DCP) converts claims about these results into executable recovery and feedback tests.\ \ Gate 1 validates that a useful improvement is achieved on a sealed evaluation. Gate 2 gives matched agents the registered starting information and observed web content while withholding the target research history. Every method that reaches the numerical target must supply a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries and a finite‑sample bound on recovery in one fresh registered episode.\ \ Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the full protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, yielding an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, passing 60‑pair null studies. Additional cases illustrate Core, recovered and audit‑incomplete decisions. A deterministic, LLM‑free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes and feedback effects across AI research.\ \ Review