THE LAB
SECTION 11
ISSUE 001
Replication Agents Become Lab Infrastructure
Projection: PaperBench directly evaluates end-to-end paper replication and finds substantial agent headroom; RE-Bench compares agents with human experts under time budgets. Institutions could run replication agents before reusing published results. The infrastructure earns trust when it surfaces missing details and produces reproducible environments that independent researchers can verify.
Why this idea is here
What the evidence establishes.
PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; RE-Bench evaluates frontier agents against humans on time-bounded machine-learning research engineering tasks. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.
Source ledger
Read the sources.
- S01PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09
- S02RE-Bench
peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09
- S03Task-completion time horizons of frontier AI models
current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10