Google DeepMind announced the pilot of double-blind AI evaluations — an approach that, in the publication's title, is described as the world's first. The project description relates to verifying closed model benchmarks in cryptographically protected environments.

According to the available description, the goal of the initiative is to increase trust in such comparative tests. Details about the trial design, participants and results in the provided source are not disclosed.

The practical significance of the project will depend on how independent the evaluations turn out and whether the proposed process can make model comparisons verifiable to external participants. For now this remains a question for the publication of data in the future.