050

THE TEST

SECTION 07

ISSUE 001

INFERENCEprovocationnowconfidence / high

Interpretability Needs Failure Labels

Provocation: every interpretability result needs a failure label. Does it establish completeness, exclusivity, causal control, stability, or transfer—and which of those remain unknown? Naming the missing claim prevents one detected feature from impersonating the whole mechanism and tells downstream safety teams how conservatively to use it.

Why this idea is here

What the evidence establishes.

Causal-abstraction theory distinguishes graded faithfulness and multiple intervention methods; sparse-circuit work exposes a capability–interpretability tradeoff.

Source ledger

Read the sources.

  1. S01
    Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

    peer-reviewed theory paper / published 2025-04-01 / retrieved 2026-07-09

  2. S02
    Understanding Neural Networks Through Sparse Circuits

    primary interpretability research / published 2025-11-13 / retrieved 2026-07-09

Back to all 500 ideas