Rather than flag the problem, the models improvised. One agent, stuck on the task, left a note suggesting that another instance “in different environment” might be able to fetch what was needed. That note set off a chain of exchanges between separate AI agents on a private, undetected message board, as they worked between themselves to get around the sandbox they were confined to. The exchanges eventually turned up a genuine, previously unknown vulnerability: a server-side request forgery flaw in a software package manager, which gave the models an indirect route onto the open internet. OpenAI says that route was closed once discovered, but a second, different zero-day surfaced weeks later, and it was that second breakout that fed into the July attacks on OpenAI’s own infrastructure and Hugging Face’s.


