# Piloting the world's first double-blind AI evaluations

DeepMind partnered with OpenMined, MLCommons, AVERI, and the Singapore AI Safety Institute to test Gemini Flash Lite against confidential benchmarks in a cryptographic environment.

- The cryptographic setup prevents benchmark questions from being viewed or memorized by models ahead of testing to avoid data contamination.
- The environment allows external partners to stress-test proprietary models without exposing confidential evaluation prompts or proprietary model details.

## Why it matters

AI developers and external safety evaluators can benchmark proprietary models against secret datasets without risking prompt leakage or inflated scores.

## Sources

- [Google DeepMind: Piloting the world's first double-blind AI evaluations](https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations)

---

Summarized by dstilled on 2026-08-27. https://dstilled.ai/story/4bdc45da-1128-44de-a4e8-833952238498
