Google launches a pilot of double-blind AI evaluations, keeping external evaluations in a cryptographic "box" to stop benchmark contamination and protect IP (Google DeepMind)
The pilot lets outside evaluators test Gemini without seeing Google’s model weights or exposing their private prompts.
Google says the setup runs evaluations inside a cryptographically secure environment, aiming to reduce benchmark contamination while protecting proprietary IP. The first test involved Gemini 2.5 Flash-Lite and a non-public evaluation set, with AVERI, OpenMined, Singapore AISI, and MLCommons named around the work. The pitch is that closed-weight models can face external safety and performance checks without either side handing over its sensitive material. Techmeme's note
Google says the setup runs evaluations inside a cryptographically secure environment, aiming to reduce benchmark contamination while protecting proprietary IP. The first test involved Gemini 2.5 Flash-Lite and a non-public evaluation set, with AVERI, OpenMined, Singapore AISI, and MLCommons named around the work. The pitch is that closed-weight models can face external safety and performance checks without either side handing over its sensitive material. Techmeme's note
score 5