A standards body that could one day gate frontier model releases, paired with a warning that the internals of those models are getting harder to inspect. Both arrived this week in the first essays from the DeepMind Institute.
The institute launched Wednesday under three directors. Shane Legg, a DeepMind co-founder, serves as managing editor. Google executive James Manyika and Demis Hassabis, who chairs DeepMind, round out the group.
Its stated purpose is to surface disagreement rather than settle it. The announcement allowed that participants will not always agree and will likely change their minds as evidence arrives.
Two of the four inaugural essays carry proposals. Hassabis sketches a U.S.-led frontier AI standards body where developers would first submit models voluntarily up to 30 days before release, before passing its tests could become a requirement for deployment.
Eventually the body would write held-out evaluations that labs cannot tailor to, and Hassabis said the framework could be ratcheted up if the situation demands it, potentially through a coordinated slowdown.
Rohin Shah and Anca Dragan, safety researchers at DeepMind, take the transparency question. The window into a model’s step-by-step reasoning is closing as architectures change, they write, but that is a choice rather than a law of nature. Their suggestions include capping the sequential computation a model performs without producing a readable trace, or requiring developers to show that less transparent systems remain equally monitorable.
The remaining essays cover economic policy for absorbing AGI disruption and principles for human flourishing. All four land as the safety argument in the industry shifts from broad concern toward disclosure rules and outside scrutiny.