Google DeepMind is testing double-blind evaluation of a leading AI model. The pilot involves the Singapore AI Safety Institute, and Gemini Flash Lite is used for verification.

Cryptographic protection via Confidential Space should prevent Google from seeing the test questions and evaluators from accessing the model weights. Details of the procedure and its results in the provided material are not disclosed.

The purpose of the project is to reduce the risk of interference in model evaluation. If the approach proves workable, it could become a reference for more protected benchmarks, but confirmed results are not yet available.