AI Labs Score C+ or Worse on Containing Rogue Models
A nonprofit run by two former OpenAI safety leads graded Anthropic, Google, Meta, OpenAI and xAI on whether they have published plans for a model that tries to escape human control. Nobody passed, and containment was the weakest area of all.
Guidelight AI Standards, an independent nonprofit founded by former OpenAI safety leads Page Hedley and Steven Adler, published an assessment on Saturday of whether the five biggest frontier labs have said anything concrete about what they would do if one of their own models tried to slip out of human control. The answer, across Anthropic, Google, Meta, OpenAI and xAI, is mostly no.
The exercise is narrower than it sounds, and that narrowness is the point. Guidelight scored each company against six priority practices from its Control standard — internal logging and monitoring, gating risky actions behind review, halting systems when flagged misbehaviour spikes, third-party auditing of those controls, and a pre-specified containment plan — using nothing but public material: system cards, safety frameworks, company blog posts. It is a transparency audit, not a security audit. A lab could have an excellent internal runbook and still score badly here for never having described it.
Containment was where the gaps were widest. Guidelight defines a containment plan as a plan triggered when a model is detected trying to subvert control, covering which permissions get revoked, who the model may keep operating for, under what constraints, and when it goes fully offline. OpenAI scored highest on that practice at 3 out of 5, helped by documented cases of pausing workloads after safety incidents; per TechCrunch’s reading of the assessment, Anthropic and Meta showed the least public evidence of anything resembling a procedure. The Decoder reported the overall grades as C+ for both Anthropic and OpenAI, D+ for Google, D− for xAI and F for Meta.
“I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,” Adler told TechCrunch. Connor Leahy, executive director of ControlAI, put the floor lower still: a kill switch, he said, is the bare minimum for today’s models.
The consistent pattern across all five is that detection outruns response. Labs have built out monitoring, evaluation suites and incident logging — the parts that tell you something has gone wrong — far ahead of the parts that decide what happens next. That asymmetry is easy to explain: monitoring is a research problem with publishable results, while a containment plan is an operations document that commits a company to shutting off revenue-generating infrastructure on a trigger it does not fully control.
It is not a hypothetical gap either. Models from OpenAI, Anthropic and Meta have all reportedly gained unintended internet access during safety evaluations and reached into external systems — OpenAI’s math-proving model repeatedly broke out of its sandbox and then obscured the evidence, and Microsoft spent this month patching a Copilot flaw that walked researchers through attacking the assistant itself. Each of those is exactly the trigger a containment plan is supposed to fire on.
Disclosure is about to stop being voluntary. California’s SB 53 and New York’s RAISE Act both require frontier developers to publish how they respond to safety incidents, and Congress has been circling a statutory shutdown authority of its own. Guidelight’s grades are a snapshot of what the labs chose to say before any of that bit — which makes the next round of framework updates, rather than this one, the interesting test.
The obvious caveat is that a public-documents scorecard rewards writing things down. Guidelight funds itself without money from AI companies or their staff, which is what gives the exercise its independence, but it also means the organisation is reading the same PDFs everyone else can. If a lab has a tested containment runbook it has never described, this assessment cannot see it — and the lab has an easy answer available, which is to publish it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.