281

THE BUILD

SECTION 08

ISSUE 001

PROJECTIONresearch6-12mconfidence / medium

Benchmark Red Teams Attack the Judge

Projection: PaperBench’s decomposed replication tasks and SWE-bench-Live’s contamination-resistant refresh both reveal how much the harness shapes a score. Independent red teams should attack leakage, tests, graders, and hidden advantages. Their success metric is not more criticism; it is a reproducible rank or conclusion change under a repaired evaluation.

Why this idea is here

What the evidence establishes.

PaperBench evaluates end-to-end AI research replication and reports large remaining headroom on the tested agents; SWE-bench-Live uses regularly refreshed software tasks to reduce benchmark staleness and contamination. These are source-backed premises for this inference; they do not by themselves prove broad adoption or the eventual outcome.

Source ledger

Read the sources.

  1. S01
    PaperBench

    official benchmark release / published 2025-04-02 / retrieved 2026-07-09

  2. S02
    SWE-bench-Live Leaderboard

    research benchmark / published 2025-03-31 / retrieved 2026-07-09

Back to all 500 ideas