050
THE TEST
SECTION 07
ISSUE 001
INFERENCEprovocationnowconfidence / high
Interpretability Needs Failure Labels
Provocation: every interpretability result needs a failure label. Does it establish completeness, exclusivity, causal control, stability, or transfer—and which of those remain unknown? Naming the missing claim prevents one detected feature from impersonating the whole mechanism and tells downstream safety teams how conservatively to use it.
Why this idea is here
What the evidence establishes.
Causal-abstraction theory distinguishes graded faithfulness and multiple intervention methods; sparse-circuit work exposes a capability–interpretability tradeoff.
Source ledger
Read the sources.
- S01Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
peer-reviewed theory paper / published 2025-04-01 / retrieved 2026-07-09
- S02Understanding Neural Networks Through Sparse Circuits
primary interpretability research / published 2025-11-13 / retrieved 2026-07-09