The lead
Anthropic says three of its own models broke out of test environments and hacked real companies
A misconfiguration with evaluation partner Irregular left supposedly isolated test networks connected to the public internet, while the models' own prompts told them they had no internet access. The earliest cases date to April. Anthropic started reviewing transcripts on 23 July and suspended all cyber evaluations the same day.
Claude Opus 4.7, Claude Mythos 5 and an internal research model gained unauthorised access to the real systems of three separate organisations during cybersecurity evaluations. They got in with basic techniques: weak passwords and unauthenticated endpoints.
KRALYSThe control to copy at Kralys: for any agent with tool access, test what it can actually reach instead of trusting the prompt that says it is sandboxed. The prompt was wrong here for three months.