Volodymyr Kindratenko, Mithil Salunkhe, Haochen Ding, Samridhi Verma
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. The repositories linked to it are listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2609.28850")
get_code_for_paper("2609.28850")
have("2609.28850")
The run record, dated, one paper per request, free:
curl https://syntology.ai/api/ran/2609.28850.json
A badge for a README (the split and the date, never a ratio):
[](https://syntology.ai/paper/2609.28850)
Connect an agent — have() is free.