Tag: AI evaluation

DeepMind locks frontier model tests in a cryptographic box

Google DeepMind is piloting the first double-blind evaluation of a proprietary frontier model to fight benchmark contamination.