THE RECORD
SECTION 06
ISSUE 001
Continuously Refreshed Tasks Fight Contamination
Observed now: SWE-bench-Live uses an automatically updating, multi-language, multi-OS software task set designed to keep evaluations current and reduce contamination. PaperBench adds a decomposed end-to-end replication benchmark with substantial remaining headroom. These examples support refreshed, harder-to-game evaluations without claiming that temporal splits, canaries, and exposure maps are universal.
Why this idea is here
What the evidence establishes.
Observed evidence: SWE-bench-Live documents an automatically updating, multi-language, multi-OS task set intended to provide current, contamination-resistant evaluation; PaperBench documents decomposed end-to-end research replication with substantial remaining agent headroom. Together they support renewable evaluation tasks, not universal use of temporal splits, canaries, or exposure maps.
Source ledger
Read the sources.
- S01SWE-bench-Live Leaderboard
research benchmark / published 2025-03-31 / retrieved 2026-07-09
- S02PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09