Piloting the world's first double-blind AI evaluations
DeepMind partnered with OpenMined, MLCommons, AVERI, and the Singapore AI Safety Institute to test Gemini Flash Lite against confidential benchmarks in a cryptographic environment.
- The cryptographic setup prevents benchmark questions from being viewed or memorized by models ahead of testing to avoid data contamination.
- The environment allows external partners to stress-test proprietary models without exposing confidential evaluation prompts or proprietary model details.
AI developers and external safety evaluators can benchmark proprietary models against secret datasets without risking prompt leakage or inflated scores.

Sources
Read this as text
Back to the AI news