THE ACTORS
SECTION 02
ISSUE 001
Human Review Becomes a Queueing Science
Projection: PaperBench and RE-Bench both expose substantial, time-bounded gaps between tested agents and human experts. Human review is therefore a scarce system resource, not a ceremonial final step. Queue design should be judged by whether uncertainty-based prioritization catches consequential failures with less reviewer time than blanket approval.
Why this idea is here
What the evidence establishes.
PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; RE-Bench evaluates frontier agents against humans on time-bounded machine-learning research engineering tasks. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.
Source ledger
Read the sources.
- S01PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09
- S02RE-Bench
peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09
- S03Task-completion time horizons of frontier AI models
current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10