THE BUILD
SECTION 08
ISSUE 001
Benchmark Red Teams Attack the Judge
Projection: PaperBench’s decomposed replication tasks and SWE-bench-Live’s contamination-resistant refresh both reveal how much the harness shapes a score. Independent red teams should attack leakage, tests, graders, and hidden advantages. Their success metric is not more criticism; it is a reproducible rank or conclusion change under a repaired evaluation.
Why this idea is here
What the evidence establishes.
PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; SWE-bench-Live uses regularly refreshed software tasks to reduce benchmark staleness and contamination. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.
Source ledger
Read the sources.
- S01PaperBench
official benchmark release / published 2025-04-02 / retrieved 2026-07-09
- S02SWE-bench-Live Leaderboard
research benchmark / published 2025-03-31 / retrieved 2026-07-09