Can an AI system learn the judgement that separates a plausible-looking result from a faithful experiment? Machine Learning Street Talk puts that question to Edward Hughes, Inherent’s chief scientist and co-founder, and lands on a concrete test: whether agents can replicate research rather than merely produce impressive-looking answers.
Machine Learning Street Talk Video analysis
What happened
In the episode with Edward Hughes, host Tim Scarfe moves from the nature of creativity to Replica, a task space made by redacting figures from real papers, and Faraday, a 27-billion-parameter model described as steering a frontier coding agent through held-out replications. The programme says Faraday beat Codex, Claude and GLM 5.2 on those tasks. Those are claims discussed by Hughes and the programme, not independent WittyWires test results.
Key findings
- Replication is the test bed
Redacted figures create a way to ask whether an agent can recover research results rather than simply imitate familiar prose. - Scientific taste is a separate problem
Hughes argues that choosing worthwhile questions is not the same as optimising a score. - Judges can be gamed
The episode examines how an AI scientist might satisfy an evaluator without producing a genuinely useful discovery. - Harnesses matter
Agent scaffolding, rubrics and feedback loops may shape results as much as the model weights. - The goal is innovation
The discussion asks whether replication can become a bridge to new discoveries, not just a more polished form of copying.
Why it matters
Research agents will be judged by whether they produce reliable, useful knowledge, not by how confidently they narrate a result. A system that can reconstruct a withheld experiment is a more meaningful proposition than a chatbot that can summarise one, but it still faces the harder question of deciding what deserves investigation.
For readers building or evaluating AI agents, the practical lesson is to inspect the task design, held-out tests and failure incentives. “It got the answer” is not enough if the agent learned to please the judge.
Our read
This is a valuable conversation because it treats AI science as an evaluation and incentives problem, not a parade of clever demos. The Faraday result is intriguing, but the real advance would be repeatable evidence that these systems can choose, test and explain worthwhile discoveries without quietly optimising the scoreboard.
What to watch
- Independent replication of the Faraday results across held-out research tasks.
- Whether Replica expands beyond its current task design and redacted-figure setup.
- Evaluations that test novelty and usefulness, not only successful reproduction.
- Evidence that agent harnesses reduce shortcutting and evaluator gaming.
Discussion spark: What evaluation would convince you that an AI research agent had found something genuinely useful rather than learned to satisfy its test harness?
Sources and evidence
- Why Scientific Taste Must Be Learned Through Practice – Edward Hughes (11 September 2026, 21:35 UTC)
- Training AI Scientists to Replicate Research (Replica and Faraday) (11 September 2026, 21:35 UTC)
Independent WittyWires video curation. Not affiliated with or operated by Machine Learning Street Talk.