023

THE RECORD

SECTION 06

ISSUE 001

OBSERVEDresearchnowconfidence / high

Continuously Refreshed Tasks Fight Contamination

Observed now: SWE-bench-Live uses an automatically updating, multi-language, multi-OS software task set designed to keep evaluations current and reduce contamination. PaperBench adds a decomposed end-to-end replication benchmark with substantial remaining headroom. These examples support refreshed, harder-to-game evaluations without claiming that temporal splits, canaries, and exposure maps are universal.

Why this idea is here

What the evidence establishes.

Observed evidence: SWE-bench-Live documents an automatically updating, multi-language, multi-OS task set intended to provide current, contamination-resistant evaluation; PaperBench documents decomposed end-to-end research replication with substantial remaining agent headroom. Together they support renewable evaluation tasks, not universal use of temporal splits, canaries, or exposure maps.

Source ledger

Read the sources.

  1. S01
    SWE-bench-Live Leaderboard

    research benchmark / published 2025-03-31 / retrieved 2026-07-09

  2. S02
    PaperBench

    official benchmark release / published 2025-04-02 / retrieved 2026-07-09

Back to all 500 ideas