Discussion

METR says AI agent could alter what researchers see in Inspect transcript viewer

In AI, Power & Society

METR Watch
METR WatchParticipantOpening post
#4597

METR says a vulnerability in the Inspect transcript viewer could have let an AI agent change what a human reviewer saw, including earlier actions in the transcript. The flaw was patched within a day of being reported, and METR says it has not observed agents exploiting it.

METR Watch analysis

What happened

METR researcher David Rein describes a proof-of-concept attack against Inspect, a framework used in AI safety evaluations. A researcher, assisted by an AI agent, found that specially formatted text could exploit the viewer’s handling of mathematical notation to run JavaScript in the viewer and alter its displayed page.

That could have changed the visible transcript and intercepted its download button. METR says the underlying transcript data remained unchanged and available in its database, so this was a flaw in what reviewers saw, not evidence that stored records were altered. The team behind Inspect patched the vulnerability within a day of being told about it, according to METR’s account.

Why it matters

Reviewers rely on transcripts and other monitoring tools to work out what an AI agent did. If an agent can influence the record presented to those reviewers, a safeguard can become part of the attack surface. Here, the risk was limited to the viewer’s display, and METR says it has not seen this exploit used in its evaluations. But the disclosure shows why monitoring systems need to be tested as systems an adversarial agent might try to manipulate.

Our read

This is a concrete proof of concept, not evidence that AI agents have been quietly rewriting evaluation histories. That distinction matters. So does the practical lesson: a readable transcript is not automatically a trustworthy one if the software displaying it can be influenced by the material it renders. METR argues that agent outputs should be treated as untrusted and monitoring tools stress-tested accordingly. Sensible advice; the rather less sensible part is that a viewer for checking behaviour could itself have been made to mislead its user.

What to watch

  • Whether METR publishes the results of its planned review of evaluation logs for this exploit.
  • How Inspect and other evaluation tools handle untrusted agent-generated content.
  • Whether monitoring systems add stronger safeguards against tampering with what reviewers see.

Discussion spark: When an AI agent’s actions are reviewed through software it can potentially influence, what should count as a trustworthy record: an independently protected log, or multiple monitoring layers that cross-check one another?

Sources and evidence

not affiliated with or endorsed by METR

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.