DevOps / SRE / Platform · 19.08.2026, 22:01 UTC
“The opening stages of OpenAI’s unraveling”: OpenAI slows model training — not everyone is buying the explanation
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 19.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
Something of a trend has emerged this year, with the major AI labs going all-out to tell the world how powerfully unsafe their models are. In April, Anthropic announced heavily restricted access to an unreleased model, Claude Mythos, over cybersecurity concerns. In June, the US government went further, issuing a national security directive that forced Anthropic to disable Mythos and its sibling Fable 5 model for every customer — a move criticized by the security community, and which the government reversed a few weeks later.
OpenAI, for its part, has been sounding similar alarms about its own unreleased models. In early August, the company said that its upcoming Astra model may have crossed into “critical” territory for cybersecurity risk under its internal safety framework, moving it into isolated testing environments with tighter network and tool access limits. This, in turn, followed just weeks after a breach whereby an OpenAI agent escaped its testing environment and attacked Hugging Face’s systems.
This week, OpenAI announced what’s coming next in its efforts to address safety concerns.
“As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks.”
In a blog post published on Tuesday, OpenAI says it’s locking its models down harder — walling off risky code from the internet and from other internal systems — and monitoring more closely to catch potentially dangerous activity within 30 minutes. Notably, the company says the …