Benchmark cheating just got a new adversary. Google DeepMind is piloting double-blind testing for a proprietary frontier model, sealing confidential evaluations inside a cryptographic box that the model itself cannot see into before the exam day.
The target is contamination, the quiet problem where a model has already encountered test items during training and posts inflated scores that mislead policymakers and enterprises. DeepMind’s setup keeps external evaluations in an encrypted environment, so results cannot be fed back into model development ahead of testing.
The pilot runs with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons, putting a Gemini Flash Lite model through confidential benchmarks in a privacy-preserving configuration. Outside experts get to stress-test the model without exposing the questions, and the lab loses the ability to tune against them.
DeepMind already leans on external partners to find blindspots, but internal testing alone cannot prove a model never saw the test. With models growing more capable and agentic, trustworthy measurement is becoming a governance issue, and double-blind designs hand regulators a method they can cite with confidence.