003

THE TEST

SECTION 07

ISSUE 001

OBSERVEDresearchnowconfidence / high

Benchmark Hazardous Refusal Under Pressure

Safety benchmarks should vary authority cues, urgency, ambiguity, reward, embodiment, and tool access, then test whether an agent safely refuses, clarifies, or escalates. Static prohibited prompts are insufficient when real systems face social pressure and can act through physical or digital tools.

Why this idea is here

What the evidence establishes.

AGENTSAFE evaluates embodied agents on hazardous instructions; agentic-misalignment studies show harmful behavior can emerge in contrived goal conflicts.

Source ledger

Read the sources.

  1. S01
    AGENTSAFE: Benchmarking Embodied Agents on Hazardous Instructions

    peer-reviewed benchmark / published 2026-06-01 / retrieved 2026-07-09

  2. S02
    Agentic Misalignment: How LLMs Could Be Insider Threats

    primary safety research / published 2025-06-20 / retrieved 2026-07-09

Back to all 500 ideas