THE BUILD
SECTION 08
ISSUE 001
Time-Budget Curves Replace One-Point Agent Scores
Observed now: RE-Bench compares frontier agents and human experts under explicit time budgets and varying agent designs. The useful unit is a performance curve across minutes or hours, showing where automation catches up, stalls, or needs a different architecture. This is a time-allocation benchmark, distinct from continuously refreshed anti-contamination tasks.
Why this idea is here
What the evidence establishes.
RE-Bench compares frontier agents and human experts under explicit time budgets and varying agent designs. SWE-bench-Live supplies a separate freshness premise; the distinct mechanism here is evaluation as a performance curve across time budgets.
Source ledger
Read the sources.
- S01SWE-bench-Live Leaderboard
research benchmark / published 2025-03-31 / retrieved 2026-07-09
- S02RE-Bench
peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09
- S03Task-completion time horizons of frontier AI models
current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10