THE HORIZON
SECTION 20
ISSUE 001
The Provenance-First AI Co-Scientist
Projection: PaperBench shows end-to-end replication is decomposable yet still difficult for agents; AlphaEvolve shows proposals can be run and scored automatically. An AI co-scientist should therefore maintain atomic claims, sources, calculations, assumptions, and contradiction logs. The provenance layer succeeds when another researcher can replay or reject each material step.
Why this idea is here
What the evidence establishes.
PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; Google DeepMind documents an evolutionary coding agent that proposes programs and scores them with automated evaluators. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.
Source ledger
Read the sources.
- S01PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09
- S02AlphaEvolve
official research release / published 2025-05-14 / retrieved 2026-07-09