THE BUILD
SECTION 08
ISSUE 001
Research Replication Becomes an Eval
Observed now: PaperBench evaluates end-to-end machine-learning paper replication through decomposed tasks and reports that tested agents remain well below the human baseline. RE-Bench separately evaluates agents against human experts on time-bounded research-engineering environments. Together they establish research reproduction as a measurable capability, without attributing every failure to a specific stage.
Why this idea is here
What the evidence establishes.
Observed evidence: PaperBench documents decomposed end-to-end machine-learning paper replication and substantial headroom below a human baseline; RE-Bench evaluates frontier agents against human experts on time-bounded machine-learning research engineering environments. Together they directly support research replication as an evaluation target, without enumerating every possible failure class.
Source ledger
Read the sources.
- S01PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09
- S02RE-Bench
peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09
- S03Task-completion time horizons of frontier AI models
current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10