044

THE BUILD

SECTION 08

ISSUE 001

OBSERVEDresearchnowconfidence / high

Research Replication Becomes an Eval

Observed now: PaperBench evaluates end-to-end machine-learning paper replication through decomposed tasks and reports that tested agents remain well below the human baseline. RE-Bench separately evaluates agents against human experts on time-bounded research-engineering environments. Together they establish research reproduction as a measurable capability, without attributing every failure to a specific stage.

Why this idea is here

What the evidence establishes.

Observed evidence: PaperBench documents decomposed end-to-end machine-learning paper replication and substantial headroom below a human baseline; RE-Bench evaluates frontier agents against human experts on time-bounded machine-learning research engineering environments. Together they directly support research replication as an evaluation target, without enumerating every possible failure class.

Source ledger

Read the sources.

  1. S01
    PaperBench

    official benchmark release / published 2025-04-02 / retrieved 2026-07-09

  2. S02
    RE-Bench

    peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09

  3. S03
    Task-completion time horizons of frontier AI models

    current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10

Back to all 500 ideas