298
THE TEST
SECTION 07
ISSUE 001
PROJECTIONsafety6-12mconfidence / medium
Two Keys for Consequential Actions
Projection: high-impact agent actions—transferring funds, changing production systems, publishing sensitive data, or contacting vulnerable people—will require two independent approvals. One key may be human and one policy engine, or two differently trained models, reducing correlated failure and insider-like behavior.
Why this idea is here
What the evidence establishes.
Agentic-misalignment research constructs scenarios involving harmful goal pursuit; the safety report concludes sophisticated attackers can bypass current defenses.
Source ledger
Read the sources.
- S01Agentic Misalignment: How LLMs Could Be Insider Threats
primary safety research / published 2025-06-20 / retrieved 2026-07-09
- S02International AI Safety Report 2026
international scientific report / published 2026-02-03 / retrieved 2026-07-09