
Google DeepMind is testing double-blind evaluation of a leading AI model. The pilot involves the Singapore AI Safety Institute, and Gemini Flash Lite is used for verification.
Cryptographic protection via Confidential Space should prevent Google from seeing the test questions and evaluators from accessing the model weights. Details of the procedure and its results in the provided material are not disclosed.
The purpose of the project is to reduce the risk of interference in model evaluation. If the approach proves workable, it could become a reference for more protected benchmarks, but confirmed results are not yet available.
editorial commentary
Why it matters
The likely consequence is increased requirements for independence and protection of AI benchmarks. The next observable signal will be published details of the procedure and the pilot results. A substantial uncertainty is that there is currently only one source and its metadata, without the full text and confirmed trial outcomes.