346

THE HORIZON

SECTION 20

ISSUE 001

PROJECTIONproduct6-12mconfidence / medium

The Provenance-First AI Co-Scientist

Projection: PaperBench shows end-to-end replication is decomposable yet still difficult for agents; AlphaEvolve shows proposals can be run and scored automatically. An AI co-scientist should therefore maintain atomic claims, sources, calculations, assumptions, and contradiction logs. The provenance layer succeeds when another researcher can replay or reject each material step.

Why this idea is here

What the evidence establishes.

PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; Google DeepMind documents an evolutionary coding agent that proposes programs and scores them with automated evaluators. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.

Source ledger

Read the sources.

  1. S01
    PaperBench

    official benchmark release / published 2025-04-02 / retrieved 2026-07-09

  2. S02
    AlphaEvolve

    official research release / published 2025-05-14 / retrieved 2026-07-09

Back to all 500 ideas