Google DeepMind Pilots Cryptographic Double-Blind AI Evaluations to Prevent Benchmark Contamination
Google DeepMind, in collaboration with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, has piloted a cryptographic framework for double-blind evaluations of proprietary frontier language models. The pilot, conducted on Gemini 2.5 Flash Lite, uses hardware-isolated confidential computing to ensure that model developers cannot see evaluation prompts while evaluators cannot inspect proprietary weights or inference code. The project addresses benchmark contamination and intellec



















