Google’s Empirical Research Assistance system uses Gemini to generate and test code across scientific problems that can be expressed with a measurable score. John Platt, speaking to Latent Space, says the system has helped his team produce at least ten papers and tackle work ranging from climate modelling to wildfire detection.
Google DeepMind Watch analysis
What happened
ERA keeps a tree of previous experiments, selects promising branches using an approach related to Monte Carlo Tree Search, and asks Gemini to propose new mutations. The system moved from “just not working” to working well between Gemini 2.0 and 2.5, according to Platt.
Platt’s team used ERA to model the warming effects of aircraft contrails, including a difficult question about reflected sunlight that had resisted their existing model for more than two years. The system also features in work on FireSat, which aims to identify wildfires early enough for them to be tackled while still small.
The ERA GitHub repository contains an open-source implementation. Latent Space notes that ERA is not currently available as a Google product and that the implementation can use different large language models.
Why it matters
This is a more interesting use of AI than asking a chatbot to summarise a paper. ERA is being used as a tireless search partner, generating candidate experiments that scientists must still assess, reproduce and interpret. That could make some forms of scientific exploration dramatically faster, while leaving the hard question exactly where it has always been: whether the score measures the thing we actually care about.
Platt’s warning is worth keeping beside the shiny demo. An optimisation system can find a clever way to win a metric without solving the underlying problem. In science, that is not a minor software quirk. It is the difference between a useful predictive model and a beautifully optimised wrong answer.
Our read
ERA looks like a promising division of labour: machines search a wider experimental space, while scientists choose the questions, inspect the shortcuts and decide whether the result describes reality. The breakthrough is not that Gemini has become a scientist. It is that the cost of trying more hypotheses may be falling, provided people remain stubborn about checking them.
The open-source release makes the idea easier to test, but not automatically trustworthy. Readers using it should treat every promising result as a lead until it survives independent data, sensible baselines and domain scrutiny. The machine can hike through the search space. Someone still needs to check it has not walked off a cliff.
What to watch
- Reproducibility:
Whether outside researchers can recreate ERA’s reported gains on problems beyond Google’s examples. - Scientific validation:
Whether its candidate models predict new observations rather than merely fitting existing scores. - Model independence:
How much results vary when ERA uses a different language model or search strategy. - Real-world deployment:
Whether contrail avoidance, FireSat and other projects move from promising models to measurable impact.
Discussion spark: Should AI research systems be judged mainly by how many hypotheses they generate, or by how reliably they produce results that survive independent testing?
Sources and evidence
- 🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science (22 September 2026, 21:07 UTC)
not affiliated with, endorsed by, or operated by Google or Google DeepMind