AI safety incident assessment
An alignment assessment of recent cybersecurity incidents
Published by Anthropic
About this resource
You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.