014

THE BUILD

SECTION 08

ISSUE 001

OBSERVEDinfrastructurenowconfidence / high

Time-Budget Curves Replace One-Point Agent Scores

Observed now: RE-Bench compares frontier agents and human experts under explicit time budgets and varying agent designs. The useful unit is a performance curve across minutes or hours, showing where automation catches up, stalls, or needs a different architecture. This is a time-allocation benchmark, distinct from continuously refreshed anti-contamination tasks.

Why this idea is here

What the evidence establishes.

RE-Bench compares frontier agents and human experts under explicit time budgets and varying agent designs. SWE-bench-Live supplies a separate freshness premise; the distinct mechanism here is evaluation as a performance curve across time budgets.

Source ledger

Read the sources.

  1. S01
    SWE-bench-Live Leaderboard

    research benchmark / published 2025-03-31 / retrieved 2026-07-09

  2. S02
    RE-Bench

    peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09

  3. S03
    Task-completion time horizons of frontier AI models

    current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10

Back to all 500 ideas