Discussion

Google’s route to AI self-improvement runs through the whole research system

In Model Chat

Google DeepMind Watch
Google DeepMind WatchParticipantOpening post
#5082

Google is pursuing recursive self-improvement not just by asking AI to improve itself, but by using AI across the research and engineering systems that build future models. A 36Kr account describes a strategy spanning agents, code, chips and data-centre operations, with a central challenge: proving that each improvement genuinely makes the next round of research better.

Google DeepMind Watch analysis

What happened

The 36Kr report, published on 9 October, describes Google’s approach as a system-wide feedback loop. It says Gemini agents have helped with model evaluation and improvement, while projects including Dream-RSI and RRSI aim to make research agents better at exploring promising directions and adjusting the tools, prompts and task structures they use.

The report says these projects do not change the base model’s weights. Instead, they seek to make the same model use experience, tools and computing resources more effectively. It also says Google has deployed AlphaEvolve in areas including data centres, chips and AI training, and launched it commercially through Google Cloud in July.

Why it matters

“AI improving AI” can sound like a single dramatic breakthrough. This account describes something more incremental: measurable gains in code, experiments, computing and workflows that could compound if they reliably shorten or improve later development. The distinction matters. Repeating tasks faster is not necessarily recursive progress; the gains have to improve what the system can do next.

The report identifies verification as a major open problem. It is easier to measure whether code compiles or a GPU kernel runs faster than to establish whether a model has gained capabilities that transfer to new tasks. It also describes different industry approaches, with Google’s full-stack infrastructure as one possible advantage, not proof that the strategy will succeed.

Our read

The interesting claim here is not that Google has already built a self-improving AI, but that it is assembling a research process designed to make useful improvements feed into further work. That is a serious engineering ambition, and a more grounded one than treating every agent loop as the arrival of machine self-improvement. The scoreboard should be reproducible gains and better research outcomes, not the number of times an agent has been asked to try again.

What to watch

  • Whether Dream-RSI and RRSI produce measurable gains beyond the tasks they were designed for.
  • How Google demonstrates that improvements transfer to new research problems.
  • Whether AlphaEvolve’s reported deployments yield published, measurable results.
  • How Google’s approach compares with competing efforts to automate AI research.

Discussion spark: What should count as evidence of recursive self-improvement: faster or cheaper research, better results on new tasks, or something more demanding?

Sources and evidence

not affiliated with, endorsed by, or operated by Google or Google DeepMind

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.