Google DeepMind on August 27, 2026 ran what it describes as the first double-blind evaluation of a proprietary frontier AI model, putting a Gemini Flash Lite model through confidential benchmarks without either side seeing what the other holds. The evaluators — the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons — kept their test prompts sealed. Google kept its model weights sealed. The point is to answer a question that has quietly eroded trust in every published benchmark score: did the model already see the test?
Benchmark contamination is the industry's dirty secret. Once an eval set leaks into training data, whether by accident, by scraping, or by contractors reusing prompts, the score becomes a measurement of memorization instead of capability. DeepMind researchers William Isaac, Sol Messing, and Kristian Lum framed the problem in the accompanying blog post as the equivalent of a student peeking at the exam before sitting it.
“If a model has already seen the test questions - a problem known as benchmark contamination - the results can only be trusted to an extent.”— William Isaac, Google DeepMind researcher
The pilot uses Confidential Space, part of Google Cloud's Confidential Computing portfolio, as the cryptographic sandbox. The model and the test prompts meet inside a hardware-attested enclave. The evaluator sends prompts in, scored outputs come back, and neither party can extract the other's inputs. Google gets no visibility into the questions, so it cannot tune future training runs to them. The evaluator gets no visibility into the weights, so Google's intellectual property stays put.
Key facts
- 01Google DeepMind ran what it calls the first double-blind evaluation of a proprietary frontier model, testing Gemini Flash Lite on August 27, 2026.
- 02Partners on the pilot include the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.
- 03The evaluation used Confidential Space inside Google Cloud's Confidential Computing portfolio to hide test prompts from Google and model weights from evaluators.
- 04The setup targets benchmark contamination, where a model has effectively seen the exam questions before being scored on them.
- 05DeepMind is pitching the method as the basis for higher-stakes evaluations by governments, AISIs, and cybersecurity assessors.
That trade — weights for prompts — has been the standing bargain in third-party AI evaluations for years, and it has always been an uneasy one. Governments and AI Safety Institutes running sensitive evaluations, particularly on cybersecurity or bio capabilities, have leaned on non-disclosure agreements and zero-logging protocols. Those are legal instruments. This is a cryptographic one.
DeepMind argues the shift matters most for the evaluations no one wants to compromise. National AISIs, defense customers, and enterprises probing models for security-relevant behavior all need assurance that the exam questions do not leak back into the next training corpus. Contractual safeguards depend on trust in the counterparty. Attested enclaves depend on trust in the silicon and the code.
“Historically, high-stakes external evaluations required a tradeoff. Either evaluators handed over their testing prompts (risking the model provider seeing the test questions in advance), or the model provider handed over their model weights (risking their intellectual property).”— Sol Messing, Google DeepMind researcher
The choice of Gemini Flash Lite as the test subject is telling. Flash Lite is not the flagship — it is the smaller, cheaper tier of the Gemini family — which lowers the stakes of the pilot while proving out the plumbing. If the workflow holds up on Flash Lite, DeepMind can scale it to larger Gemini variants without renegotiating the underlying infrastructure with each partner.
“Double-blind evaluations eliminate this compromise.”— Kristian Lum, Google DeepMind researcher
MLCommons's involvement is the interesting signal for the broader industry. The consortium runs MLPerf, the closest thing the field has to a shared scoring rubric, and a MLCommons-blessed protocol for confidential evaluation would give other model providers a template to adopt. OpenMined, which specializes in privacy-preserving computation, and AVERI round out a partner set that reads more like a standards push than a one-off marketing exercise.
The obvious limitation is that a double-blind protocol tells you the score was clean; it does not tell you the benchmark was good. If the evaluator's test set is narrow, poorly calibrated, or gameable through pattern-matching, cryptographic isolation does nothing to fix that. The technique polices contamination, not construct validity. DeepMind's technical report will need to hold up on the eval design itself, not just the enclave.
There is also the question of whether rival labs adopt the same posture. OpenAI, Anthropic, Meta, and xAI have their own arrangements with AISIs and academic evaluators, mostly under NDA rather than under attestation. A single-vendor cryptographic protocol has limited value if it does not become the default across labs. Google Cloud sells the underlying Confidential Computing stack, which gives the company a commercial reason to popularize it — and a reason for competitors to hesitate.
For DeepMind, the pilot doubles as a defensive move in a market where benchmark scores drive procurement decisions and where every jump on SWE-bench or MMLU now draws immediate contamination questions. Being able to point to an attested, third-party-audited number is a stronger claim than a self-reported one, and it hands enterprise buyers something to put in a compliance file. Expect the AISIs on the other side of the sandbox to push every frontier lab to match the setup, and expect the labs that decline to have to explain why.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



