019
THE TEST
SECTION 07
ISSUE 001
OBSERVEDsafetynowconfidence / high
Safeguards Trigger on Capability
Trigger safeguards on demonstrated capability, not a model’s brand or parameter count. Independently replicated tests for cyber operations, biological uplift, deception, and autonomous persistence should activate stronger controls, with conservative uncertainty and frequent threshold revision. The hard part is keeping the test ahead of new elicitation methods.
Why this idea is here
What the evidence establishes.
The 2026 safety report describes capability advances and additional safeguards; agent benchmarks show safety properties vary substantially across systems and contexts.
Source ledger
Read the sources.
- S01International AI Safety Report 2026
international scientific report / published 2026-02-03 / retrieved 2026-07-09
- S02AGENTSAFE: Benchmarking Embodied Agents on Hazardous Instructions
peer-reviewed benchmark / published 2026-06-01 / retrieved 2026-07-09