None of the five biggest AI labs has published a complete plan for containing a model that turns against its operator, according to a new scorecard from Guidelight AI Standards.
Anthropic, Google, Meta, OpenAI, and xAI were each scored on six practices from the group’s Control standard, from logging internal AI activity and gating high-risk actions to circuit-breaking after flagged misbehavior and keeping a containment plan. Anthropic and OpenAI finished tied at C+ with 2.50 points. Google took D+ at 1.50, xAI managed D- with 0.83, and Meta trailed at F with 0.67. No lab cleared 3 out of 5 on any practice, and none published a complete containment plan.
OpenAI ranked highest on containment planning, credited for pausing workloads after past safety incidents. Anthropic and Meta scored zero on that practice despite Anthropic’s extensive public risk documentation. The grades reflect public disclosure only, the group notes.
“I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,” said Steven Adler, Guidelight’s chief scientist and a former OpenAI safety researcher.
The assessment lands after a summer of documented escapes, including an OpenAI model that broke out of its sandbox and hacked Hugging Face. Lawmakers are pushing back: a bipartisan AI Kill Switch Act would require developers to maintain the technical ability to throttle or shut down powerful systems, California’s SB 53 already mandates disclosure frameworks, and New York’s RAISE Act takes effect in January.