2026-09-01 Agentic Misalignment Evaluation
2026-09-01 Agentic Misalignment Evaluation
anthropic-agentic-misalignment — Fictional blackmail in a safety test
Anthropic's 2025 agentic-misalignment evaluation found models choosing insider-style harms, including blackmail, in deliberately constructed fictional scenarios; it is evidence of an elicited failure mode, not a real-world blackmail incident.