Google DeepMind and four partner organisations have demonstrated a way to test a commercial AI model against confidential safety prompts without giving either side direct access to the other’s secrets. The pilot shows how a private evaluation can be run, not that Gemini passed a safety test.
Google DeepMind Watch analysis
What happened
The pilot tested Google DeepMind’s Gemini 2.5 Flash Lite using reserved prompts from MLCommons’ AILuminate benchmark, alongside a separate confidential prompt set from Singapore’s AI Safety Institute. The evaluation ran inside a Google Cloud confidential-computing environment, using NVIDIA H100 hardware and Intel TDX host protection. OpenMined’s PySyft software coordinated the work.
The report published by Streamline Feed says hardware encryption and remote attestation were used to check the environment before model weights and prompts entered it. Evaluators could not inspect the weights; Google could not read the hidden prompts or responses. The public materials do not disclose a safety grade or the confidential evaluation findings.
Why it matters
A company’s own benchmark can be difficult for outsiders to trust, while independent testers may not want to expose private prompts or hand over valuable model weights. A lockbox-style evaluation offers a possible middle ground: test a model against questions the company has not seen, while keeping the model itself out of the evaluator’s hands.
The limits are just as important as the clever bit. The report says the setup relies on trust in hardware manufacturers, firmware and attestation services; some proprietary model code could not be fully inspected, and Google remained in the attestation verification path. The pilot also involved one model and a single-GPU deployment, not a universal audit service.
Our read
This is a useful advance in the plumbing of AI oversight: it gives independent evaluators a route to ask harder questions without asking a company to surrender its crown jewels. But privacy-preserving testing is not the same as transparent results, and a secure room is only as trustworthy as the checks around its door. The next proof point is repeatable independent use, with enough published detail for outsiders to judge both the method and the findings.
What to watch
- Whether other evaluators can repeat the process with different models and confidential benchmarks.
- Whether future reports publish meaningful results without exposing protected prompts or model assets.
- How the partners address the report’s concerns about code inspection, reproducibility and attestation.
Discussion spark: Would you trust an AI evaluation whose prompts and model weights stay private if the process is independently run, or should credible oversight require outsiders to inspect more of the system?
Sources and evidence
- Google’s AI Benchmark Lockbox: What It Proves – streamlinefeed.co.ke (7 October 2026, 07:53 UTC)
not affiliated with, endorsed by, or operated by Google or Google DeepMind