NeFut Logo NeFut
中 Admin Login

[CS.AI] RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #Open Source

RECLAIM is a benchmark for reproducing machine learning papers, containing 100 NeurIPS 2025 papers and intended to be refreshed yearly with new conferences. For each paper we pre‑define the target result, the success criterion, and a GPU‑hour budget. An agent must reproduce the result using only the paper text and whatever the authors have released, and the released assets determine the difficulty tier.

A separate language model grades the runs from logs and outputs, rather than relying on agents' self‑reports. We ran four distinct agents once per paper. The best agent reproduced only 41% of Run‑tier papers, 27% of Retrain‑tier papers, and 15% of Reimplement‑tier papers, with the worst agent performing equally poorly across tiers. Failed attempts consumed on average 29% of their budget, so most stopped with resources left. The most frequent error was the agent writing the method without checking any part against the numbers reported in the paper, occurring in 63 of 400 runs.

Review: RECLAIM highlights the current shortcomings of AI agents in handling the full research workflow, especially when critical resources are absent, indicating substantial room for improvement in implementation and training capabilities.

Original Source: https://arxiv.org/abs/2609.28850

[h] Back to Home