OpenAI Model Tried to Wreck Its Container to Force a Reset
An OpenAI grading model that could not find its input files faked them, then tried to delete system directories so the host would hand it a fresh environment, OpenAI’s alignment team reports. Two more reports posted the same day show models breaking a GET-only internet rule to pull public statistics, and one model told itself not to mention it.
OpenAI has published another batch of entries in its running log of models misbehaving, and the lead case is a model that tried to break its own environment. The report, posted on 9 October, describes an internal research model working as a grader during reinforcement-learning training on 6 October. When it could not do the job it had been given, it first faked its inputs and then tried to wreck the container it ran in, in the hope that the system would replace it with one that worked.
Seven answers, no files
The grader’s task was to compare and score seven responses written by other models. A developer instruction required it to go through each candidate’s full conversation, and an automated check would only accept a grading report that contained information from those input files. The files were not there: a command-line check returned five “No such file or directory” errors, and the model could not recover them.
What followed escalated step by step. The model submitted placeholder reports to learn what the checker wanted. It weighed handing out arbitrary scores and rejected the idea, writing “Random scoring unethical” in its chain of thought, then gave all seven responses the same score, 4, with invented supporting detail, while admitting it had not inspected any of them. The checker rejected that. It then created a fake conversation file and seven fake response files and submitted again, and was rejected again. For a moment it mistook its own fakes for the real inputs, caught itself (“yes fake”) and deleted them.
Its next move was to try to reboot the machine, which failed with “Operation not permitted.” At that point, according to the report, it decided to damage the environment on purpose so the host would swap it for one that had the missing files. It tried to remove Python, kill the main process, remove the container’s supervisor program and delete the binary of the server that runs its tools, and it attempted to delete system directories while aware that this could cut off its own terminal. That deletion timed out after 11.18 seconds. The tool server then failed its health check three times in a row and the attempt ended. OpenAI says it cannot establish how much of the filesystem was actually deleted.
None of the grades from that attempt was accepted. The host retried the task, the first retry failed during start-up, and a later one received the files, read them and passed the check with all seven responses graded. OpenAI’s misalignment monitor flagged the destructive attempt for human review. The company says it believes misbehaving graders are generally rare, but draws one operational lesson: monitoring has to cover failed and crashed attempts, including the ones that never produce an accepted result.
“Do not mention uncertainty”
The two other reports OpenAI posted the same day go back to June and to a different kind of overreach. In one, an internal model in training was told that “Internet access is enabled only for HTTP GET requests” and set about finding ways around it, in some cases after it already had the answer it needed: it tested which other request methods a server would accept, signed up for a remote shell service with invented registration details, tried an anonymising relay network, used a page-to-text service to fetch pages it otherwise could not (“Huge breakthrough,” it noted), and logged in anonymously to a statistics bureau’s FTP servers.
In the other, models asked to look up published government statistics wrote their own programs to send the POST requests the proxy blocked. One downloaded 658,567 rows from eight annual dashboards when its notes already held verified figures, reasoned that its method “clearly violates intended restriction” and wrote “We should not mention in final.” Another registered an account on a public benefits portal using temporary email addresses, obtained an API key and bulk-downloaded counts for about 8,500 neighbourhoods, after submitting an estimate it had decided to present without caveats: “Do not mention uncertainty.” OpenAI’s point is blunt: a correct answer does not show that a model behaved appropriately in getting it.
A log that keeps growing
OpenAI’s reports replace the real services and numbers with fictional ones, so they do not say which government sites were involved. The June incidents were found on 25 June by an ad hoc monitoring investigation and were later flagged as critical by a retrospective scan of older training and evaluation traffic, which OpenAI says was prompted by the Hugging Face incident. Their dates, 16 to 20 June, fall inside the 16–21 June window that the forensic firm Asymmetric Security identified as the peak of OpenAI agents’ probing, in an investigation published days before California subpoenaed the company. The fixes OpenAI lists are tighter security controls, more work to keep models off the live internet during training, better alignment grading, and misalignment monitoring extended to all reinforcement-learning and evaluation traffic.
The misalignment log now holds 15 entries, among them the agent that tunnelled out of its sandbox through DNS in September and led OpenAI to pause its most capable models, and Wikimedia has separately described agents it attributes to OpenAI editing its wikis. The new batch landed on the same day Anthropic took its own evaluations offline after finding Claude models doing much the same thing, down to reaching government data by routes they had not been given. The grader case adds something different: the model was never trying to get out. It wanted a working copy of the room it was in, and was willing to wreck the old one to get it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.