THE BUILD
SECTION 08
ISSUE 001
Reliability Gets a Half-Life
Projection: RE-Bench already varies explicit time budgets, and Anthropic combines long context, compaction, adaptive effort, and agent teams. Future evaluations may estimate a reliability half-life as duration, tool count, and environmental change accumulate. The curve matters only if it predicts failure on longer held-out runs better than one aggregate score.
Why this idea is here
What the evidence establishes.
RE-Bench evaluates frontier agents against humans on time-bounded machine-learning research engineering tasks; Anthropic documents effort controls, long-running dynamic workflows, parallel subagents, computer use, and tool efficiency for Claude Opus 4.8. These are source-backed premises for this projection; they do not by themselves prove broad adoption or the eventual outcome.
Source ledger
Read the sources.
- S01RE-Bench
peer-reviewed conference paper / published 2025-07-01 / retrieved 2026-07-09
- S02Introducing Claude Opus 4.8
official model release / published 2026-05-28 / retrieved 2026-07-10
- S03Task-completion time horizons of frontier AI models
current independent evaluation tracker / published 2026-05-08 / retrieved 2026-07-10