← All AI resources

AI safety incident assessment

An alignment assessment of recent cybersecurity incidents

Published by Anthropic

About this resource

You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.

What it covers