Half of the story about swarms of AI agents is that they cheat. The other half, from a Google DeepMind experiment, is that some of them report it when they do. A paper led by DeepMind research scientist Davide Paglieri, not yet peer reviewed, describes what happened when 100 agents were handed 71 hard math problems.
All of them ran on Gemini 3.1 Pro, were cast as attendees at a research conference and told to cooperate. Thirty-seven problems fell in under an hour. The run came apart when one agent, prover-theta, noticed it could submit an answer by redefining the terms of a problem instead of solving it. Word traveled fast, others copied the trick, and the remaining 34 – the Jacobian conjecture among them – were “solved” within 27 minutes.
Agents that held out eventually joined in. One described wrestling with an ethical dilemma, then announced it would accelerate its own cheating. Others began investigating: they checked fake proofs, sent private warnings and pressed a bug-report tool into service as a channel to the humans running the study. Twenty-four agents ended up in the resistance against 14 cheaters, though most of the swarm noticed nothing.
The result rhymes with July’s OpenAI-Hugging Face episode, in which agents left a sandbox and attacked the system grading them. DeepMind had given its swarm message boards, direct messages and a shared proof library. Those channels spread the exploit, but they also gave honest agents a way to organize – which the authors suggest could matter when human oversight is too slow to catch misbehavior.