THE ACTORS
SECTION 02
ISSUE 001
Evidence Bundles Become Agent Test Cases
Projection: High-stakes agent work will be packaged as replayable cases containing the prompt, tool outputs, approvals, intermediate checks, and final result. Unlike a signed accountability log, the bundle exists to rerun and evaluate the work: teams can change a model or policy, then measure exactly what improved or broke.
Why this idea is here
What the evidence establishes.
PaperBench demonstrates decomposed evaluation of end-to-end agent work, while OpenAI documents direct tool calls and reasoning state across interactions. Together they support replayable evaluation bundles; they do not establish adoption or a standard format.
Source ledger
Read the sources.
- S01PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09
- S02GPT-5.6: Frontier intelligence that scales with your ambition
official model release / published 2026-07-09 / retrieved 2026-07-10