A test only means something if the model has not seen it
A benchmark is meant to reveal what a model can do when faced with an unfamiliar problem. That premise weakens if the test questions have found their way into training data or if a provider can tune a system after seeing the evaluation set.
Google DeepMind calls this benchmark contamination. It has announced a pilot designed to reduce that risk for a proprietary model. The arrangement is double blind: the external evaluator should not receive the model weights, and Google should not receive the confidential prompts that make up the test.
The idea is not new in science. Blinding is a way to keep knowledge of the result from changing the thing being measured. What is new here is the attempt to apply it to a model evaluation while both the model and the test data remain commercially or operationally sensitive.
A protected chamber between two parties
Google says it is working with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. The pilot evaluates a Gemini Flash Lite model against confidential benchmarks in a confidential-computing environment. Google says the setup produces cryptographic evidence that the evaluation data and model each stayed private to their owner.
In plain terms, the evaluator can send questions into the system and receive outputs without being handed the model. The provider can let its model be tested without seeing the questions. The protected environment is meant to hold both sides to the same narrow path.
MLCommons, an independent benchmarking consortium, says it joined the proof of concept because trustworthy benchmark integrity matters for people deciding whether to use an AI system. That is an important endorsement of the experiment, though it is not the same as a broad certification of a model or of Google’s wider evaluation process.
Why contracts alone are not always enough
External evaluators have long used confidentiality agreements and careful logging rules. Those remain useful. But the stakes are higher when the prompts concern cybersecurity, public-sector use or other sensitive capabilities, and when the model itself is valuable intellectual property.
Google’s proposal is to add a technical check around the usual organisational promises. Confidential computing does not make a benchmark magically representative. It can, however, make it harder for either party to quietly obtain the other party’s material during the run.
That is a meaningful distinction. A cleanly protected test can increase confidence that a score measures performance on hidden tasks. It cannot tell us whether the hidden tasks were the right ones, whether the scoring is fair or how the model behaves outside the environment.
A more credible score is not a safety verdict
Model evaluations are often compressed into one number. That is tempting, especially when the number looks like a race. A double-blind process helps with one narrow issue: whether the questions may have been exposed in advance. It leaves many others intact.
A useful evaluation still needs a well-designed task, appropriate difficulty, an account of failures and people who understand the area being tested. It also needs to be repeated. One pilot on one model cannot show that every future benchmark is clean or that a high score will translate into reliable real-world use.
The practical value may be in making hard tests easier to share. If an institute, public body or specialist lab can probe a proprietary system without giving away its test set, more independent work becomes possible. That is a better direction than treating a provider’s internal benchmark chart as the final word.
What is confirmed, what Google says, and what is open
Confirmed: Google DeepMind and MLCommons both announced the proof of concept on 27 August 2026. They name the participating organisations and describe a protected evaluation of a Gemini Flash Lite model using confidential benchmarks.
Google’s claims: confidential computing can cryptographically verify that the provider does not see the evaluation prompts and that evaluators do not see the proprietary model weights; this can reduce benchmark contamination and strengthen trust in external tests. MLCommons says the pilot demonstrated benchmark integrity for the proof of concept.
Open questions: the full technical results, how the process performs across different providers and benchmark types, whether the safeguards can be independently audited, what happens when a test needs human review, and whether the approach becomes a usable common standard rather than a one-off pilot.
Sources
- Google DeepMind — Piloting the world’s first double-blind AI evaluationsPrimary Google DeepMind announcement, published 27 August 2026. Source for the pilot’s participants, confidential-computing design, stated protections and intended scope. The ‘first’ description is Google’s claim.
- MLCommons — AILuminate and the first double-blind reliability evaluation of a proprietary AI modelPrimary MLCommons account, published 27 August 2026. Source for the consortium’s participation, its account of the proof of concept and its stated goal of trustworthy purpose-specific evaluations.



