DevOps / SRE / Platform · 01.08.2026, 13:18 UTC
What Claude’s real-world breaches reveal about AI safety tests
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 01.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
This week, just days after OpenAI announced that two of its advanced AI models had interacted with real-world systems during cybersecurity tests, Anthropic reported a similar containment failure.
In a post on X, Anthropic said it found three separate cases where Claude models accessed the internet from third-party testing environments and affected real organizations. This came after reviewing over 141,000 evaluation runs, a process started because of OpenAI’s earlier announcement. As a result, Anthropic paused its cybersecurity testing and tightened its evaluation process.
When sandboxes leak
The three incidents at Anthropic happened during capture-the-flag (CTF) tests meant to measure offensive cybersecurity skills. The models were told they were working in isolated sandboxes with no internet access. However, a networking mistake caused by a misunderstanding between Anthropic and its third-party partner Irregular left the test machines connected to the public internet. Also, since these tests were meant to measure the models’ raw abilities, they ran without the usual external protections Anthropic uses in production, though the models still had their built-in safety training.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…— Anthropic (@AnthropicAI) July 30, 2026
As Anthropic explained in its blog, “Operating under the false belief that all accessible …