003
THE TEST
SECTION 07
ISSUE 001
OBSERVEDresearchnowconfidence / high
Benchmark Hazardous Refusal Under Pressure
Safety benchmarks should vary authority cues, urgency, ambiguity, reward, embodiment, and tool access, then test whether an agent safely refuses, clarifies, or escalates. Static prohibited prompts are insufficient when real systems face social pressure and can act through physical or digital tools.
Why this idea is here
What the evidence establishes.
AGENTSAFE evaluates embodied agents on hazardous instructions; agentic-misalignment studies show harmful behavior can emerge in contrived goal conflicts.
Source ledger
Read the sources.
- S01AGENTSAFE: Benchmarking Embodied Agents on Hazardous Instructions
peer-reviewed benchmark / published 2026-06-01 / retrieved 2026-07-09
- S02Agentic Misalignment: How LLMs Could Be Insider Threats
primary safety research / published 2025-06-20 / retrieved 2026-07-09