019

THE TEST

SECTION 07

ISSUE 001

OBSERVEDsafetynowconfidence / high

Safeguards Trigger on Capability

Trigger safeguards on demonstrated capability, not a model’s brand or parameter count. Independently replicated tests for cyber operations, biological uplift, deception, and autonomous persistence should activate stronger controls, with conservative uncertainty and frequent threshold revision. The hard part is keeping the test ahead of new elicitation methods.

Why this idea is here

What the evidence establishes.

The 2026 safety report describes capability advances and additional safeguards; agent benchmarks show safety properties vary substantially across systems and contexts.

Source ledger

Read the sources.

  1. S01
    International AI Safety Report 2026

    international scientific report / published 2026-02-03 / retrieved 2026-07-09

  2. S02
    AGENTSAFE: Benchmarking Embodied Agents on Hazardous Instructions

    peer-reviewed benchmark / published 2026-06-01 / retrieved 2026-07-09

Back to all 500 ideas