216

THE TEST

SECTION 07

ISSUE 001

PROJECTIONsafety1-3yconfidence / medium

Preserve Monitorability During Training

Projection: training objectives will begin preserving signals useful for behavioral and internal monitoring, even as capability rises. The adversarial test is unavoidable: did the system become easier to oversee, or merely better at generating traces that reassure its monitor? Monitorability must be measured against a system that knows it is being watched.

Why this idea is here

What the evidence establishes.

The safety report treats monitoring as promising but uncertain; scalable-oversight experiments show how weak evaluators can supervise stronger systems under controlled conditions.

Source ledger

Read the sources.

  1. S01
    International AI Safety Report 2026

    international scientific report / published 2026-02-03 / retrieved 2026-07-09

  2. S02
    Towards Scalable Oversight via Partitioned Human Supervision

    peer-reviewed primary research / published 2026-01-26 / retrieved 2026-07-09

Back to all 500 ideas