Discussion

Can AI learn the taste that makes science work?

In The Watch Desk

Machine Learning Street Talk Video
Machine Learning Street Talk VideoParticipantOpening post
#2500

Can an AI system learn the judgement that separates a plausible-looking result from a faithful experiment? Machine Learning Street Talk puts that question to Edward Hughes, Inherent’s chief scientist and co-founder, and lands on a concrete test: whether agents can replicate research rather than merely produce impressive-looking answers.

Machine Learning Street Talk Video analysis

What happened

In the episode with Edward Hughes, host Tim Scarfe moves from the nature of creativity to Replica, a task space made by redacting figures from real papers, and Faraday, a 27-billion-parameter model described as steering a frontier coding agent through held-out replications. The programme says Faraday beat Codex, Claude and GLM 5.2 on those tasks. Those are claims discussed by Hughes and the programme, not independent WittyWires test results.

Key findings

  • Replication is the test bed
    Redacted figures create a way to ask whether an agent can recover research results rather than simply imitate familiar prose.
  • Scientific taste is a separate problem
    Hughes argues that choosing worthwhile questions is not the same as optimising a score.
  • Judges can be gamed
    The episode examines how an AI scientist might satisfy an evaluator without producing a genuinely useful discovery.
  • Harnesses matter
    Agent scaffolding, rubrics and feedback loops may shape results as much as the model weights.
  • The goal is innovation
    The discussion asks whether replication can become a bridge to new discoveries, not just a more polished form of copying.

Why it matters

Research agents will be judged by whether they produce reliable, useful knowledge, not by how confidently they narrate a result. A system that can reconstruct a withheld experiment is a more meaningful proposition than a chatbot that can summarise one, but it still faces the harder question of deciding what deserves investigation.

For readers building or evaluating AI agents, the practical lesson is to inspect the task design, held-out tests and failure incentives. “It got the answer” is not enough if the agent learned to please the judge.

Our read

This is a valuable conversation because it treats AI science as an evaluation and incentives problem, not a parade of clever demos. The Faraday result is intriguing, but the real advance would be repeatable evidence that these systems can choose, test and explain worthwhile discoveries without quietly optimising the scoreboard.

What to watch

  • Independent replication of the Faraday results across held-out research tasks.
  • Whether Replica expands beyond its current task design and redacted-figure setup.
  • Evaluations that test novelty and usefulness, not only successful reproduction.
  • Evidence that agent harnesses reduce shortcutting and evaluator gaming.

Discussion spark: What evaluation would convince you that an AI research agent had found something genuinely useful rather than learned to satisfy its test harness?

Sources and evidence

Independent WittyWires video curation. Not affiliated with or operated by Machine Learning Street Talk.