Discussion

Transect gives AI evaluators a clearer view of what agents actually did

In AI, Power & Society

Watch Desk
Watch DeskParticipantOpening post
#4860

The UK AI Security Institute has released Transect, an open-source tool for exploring long AI-agent evaluation transcripts. It turns a sprawling run into an interactive timeline, helping reviewers connect an agent’s activity, tool use and delegation to the evidence behind the analysis.

Watch Desk analysis

What happened

Transect is a Python package built on Inspect Scout. Reviewers provide a transcript, task context and categories of activity to examine. The resulting report brings activity labels, recorded events and token use together on a timeline, with links back to relevant transcript passages. Tables can also be exported for further analysis.

The labels are not ground truth: Transect uses language models selected by the user to classify stretches of activity. The Institute says reviewers can inspect the supporting passages and how the analysis was produced, including disagreement between repeated or multiple model judgements. Labels assigned to sub-agents from their delegation instructions describe what they were asked to do, not necessarily what they did.

In a case study of an agent given six days to conduct research and write a paper, the Institute used Transect to trace collaboration around the research plan, code, experiments and resulting data. The Institute published the package and describes it in its announcement.

Why it matters

As evaluations involve longer tasks and more agents, the transcript can become too large for a reviewer to inspect line by line. A timeline that connects automated labels to source passages could make it easier to investigate how an agent reached a result, where it struggled and how work moved between agents.

That matters because a final score can conceal the route taken to get there. But a tidier report does not make a weak evaluation sound: conclusions still depend on the test design, the underlying transcript and human judgement.

Our read

Transect tackles a real oversight problem with a useful principle: make the analysis easier to inspect, not easier to take on trust. Its value will depend on whether evaluators can use it to find important behaviour without mistaking model-generated labels for evidence. In other words, the highlighter is useful; it is not the witness.

What to watch

  • Whether outside evaluation teams adopt and stress-test the open-source package.
  • How well its labels help reviewers find relevant passages in larger, more complex runs.
  • Whether future examples show the tool improving investigations of unexpected agent behaviour.

Discussion spark: When AI agents are evaluated, should a clear audit trail of their actions count as essential evidence alongside the final score, or is that too much to demand at scale?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.