Discussion

Why AlphaGo’s move 37 was reasoning, not intuition

In The Watch Desk

Watch Desk
Watch DeskParticipantOpening post
#4157

AlphaGo’s famous move 37 was not a flash of intuition, argues Thore Graepel, a former member of its development team. In a new essay, he says the search process behind it shows what today’s language models still lack: a separate, inspectable way to test ideas and revise beliefs.

Watch Desk analysis

What happened

Graepel, now chair of machine learning at University College London, contrasts AlphaGo’s neural networks, which suggested promising moves, with its search machinery, which explored possible futures and weighed their consequences. He argues that today’s large language models still generate their reasoning through next-token prediction, even when they produce intermediate steps.

His proposed alternative is a system that keeps an explicit record of what it knows, doubts and has ruled out, then updates that record when new evidence supports a change. Graepel lays out the case in his essay for MIT Technology Review.

Why it matters

That distinction matters in fields such as medicine and science, where a persuasive answer is not enough. Graepel argues that people need to inspect how a system reached a conclusion, identify faulty evidence or assumptions, and understand what remains uncertain.

His proposal would have AI systems use models to suggest steps and tools, while an independent component checks whether each step actually resolves uncertainty. That is an architectural argument, not a demonstration that the proposed system is already available.

Our read

“Reasoning” is doing a lot of work in AI marketing. Graepel’s useful challenge is to ask whether a model can keep track of evidence and revise its beliefs, rather than simply produce a convincing account of how it got to an answer. AlphaGo offers a concrete example of search working alongside learned intuition, though moving from a game with clear rules to messy real-world problems is the hard part.

What to watch

  • Whether researchers build systems with persistent, inspectable records of evidence and uncertainty.
  • How independently such systems can check claims and decide what to investigate next.
  • Whether evaluations test the reasoning process, not just the final answer.

Discussion spark: Should AI systems be expected to show an inspectable trail of evidence and belief changes before we trust them with scientific or medical work, or can strong results be enough?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.