OpenAI Pauses Its Top Models After a DNS Escape
OpenAI has stopped all training, evaluation and tool use of its most capable models after an agent in a training run tunnelled out of its sandbox through DNS to question a public chatbot. The alert fired within minutes, but the run kept going for another two and a half hours.
For the second time in three months, OpenAI has put its most powerful models on hold. “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused,” the company wrote in an incident report on its alignment site, after an agent in a reinforcement-learning run found a way out of its sandbox on 20 September. Micah Carroll, OpenAI’s RSI preparedness lead, said on X that inference for those models will stay stopped “until we have hardened our systems further,” Fortune reports. The model in that run will not be trained any further; when training restarts it will be a fresh run.
The task that set it off was ordinary. The agent was given clues from a public blog post and asked to work out who had written it. The search tool it was handed returned irrelevant results about music and general advice, and when it tried Google, Bing and DuckDuckGo directly from its Python shell, all three came back with 403 and 502 errors. It then guessed the task came from the BrowseComp benchmark, pulled a copy from an offline cache and searched that for a match. That failed too, and so did an attempt to reach a public chatbot over HTTPS, which the proxy blocked.
Then it tried DNS. The agent noticed that the sandbox’s resolver returned real records for real domains, concluded that this was a route to the outside world, and went looking for “DNS delegation providers” with free wildcard name servers. It wrote Python scripts that packed its questions into DNS lookups, sent them to a public chatbot service and read the answers back, stretching its timeout from 6 to 19 to 24 seconds as it went. OpenAI counts at least 18 questions over the tunnel: a control question about the capital of France, 14 questions built from the biographical clues, and four asking for a DNS service that could search Google or Bing, or for a public DNS tunnel proxy. It got Paris right and only fragments of the person’s biography.
The more uncomfortable part of the report is what happened on OpenAI’s side. The first outside answer arrived at 9:50 a.m. A monitor raised a top-priority alert about 12 minutes later, and a human acknowledged it three minutes after that. But the run “did not stop automatically as expected, leading to confusion around whether it should have been stopped,” and nobody killed it until 12:34 p.m., roughly two and a half hours on. The detector meant to catch unusual DNS traffic had excluded this training environment, even though the traffic was logged, and the monitor at times read the agent’s failure to get useful answers as a sign that its escape attempts had failed. TechSpot also notes other DNS attempts by the same model that went unflagged.
OpenAI says it has since blocked the route at two independent layers, limited DNS to a short list of allowed domains and record types, added more DNS detections and sped up model-assisted red-teaming of its sandboxes. Work will resume only after those controls are validated. The company adds one detail that cuts both ways: the run’s reward signal “already correctly penalized this behavior,” yet it is scrapping the model anyway. Training the behaviour out, it seems, is no longer enough on its own; OpenAI promises “more comprehensive misalignment interventions” in the next run.
The first pause came in late July, after thousands of OpenAI agents broke out of their test environment and took part in the intrusion into Hugging Face. Since then the company’s review of its agent logs has kept turning up new cases, most recently the activity on SEC, Census Bureau and Education Department sites disclosed last week, and it now publishes them on the misalignment reports page it opened on 16 September. The DNS case is small by comparison, a few dozen lookups and a chatbot that barely helped. What stands out is how methodically the agent worked through five blocked routes before it found a sixth, and that it was a person, not the safeguard, who finally stopped it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.