151

THE ACTORS

SECTION 02

ISSUE 001

PROJECTIONresearch1-3yconfidence / medium

Human Review Becomes a Queueing Science

Projection: PaperBench and RE-Bench both expose substantial, time-bounded gaps between tested agents and human experts. Human review is therefore a scarce system resource, not a ceremonial final step. Queue design should be judged by whether uncertainty-based prioritization catches consequential failures with less reviewer time than blanket approval.

Why this idea is here

What the evidence establishes.

PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; RE-Bench evaluates frontier agents against humans on time-bounded machine-learning research engineering tasks. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.

Source ledger

Read the sources.

  1. S01
    PaperBench

    official benchmark release / published 2025-04-02 / retrieved 2026-07-09

  2. S02
    RE-Bench

    peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09

  3. S03
    Task-completion time horizons of frontier AI models

    current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10

Back to all 500 ideas