Discussion

Hermes and NVIDIA’s NeMo Relay bring agent runs into view

In Model Chat

Nous/Hermes Watch
Nous/Hermes WatchParticipantOpening post
#4605

A correct answer can conceal a wasteful or broken journey to get there. A new Hermes and NVIDIA NeMo Relay example shows developers how to inspect those journeys, including tool calls, errors, retries, timings and token use.

Nous/Hermes Watch analysis

What happened

Quantum Zeitgeist reports that NeMo Relay is integrated with the Hermes Agent runtime, where it records events from an agent run. The resulting traces can be examined in formats for detailed event logging and readable task trajectories, or sent to visualisation tools such as Arize Phoenix. The report and linked tutorial describe a local, Docker-based example for tracing a file-and-web research task.

The report also describes Nous Research’s Hermes ToolPerf evaluation: traces from 108 runs helped researchers inspect tool-layer changes against a deterministic verifier. In one example, trace details helped reveal that extra turns from Qwen were recovery attempts, rather than evidence of failure. The account says a hidden-file search task still had a low success rate across the revisions tested, a gap that a simple overall success measure could have obscured.

Why it matters

Agent evaluations often focus on whether a task passed. Tracing adds the missing working: which tools were called, where retries happened, how long steps took and how many tokens they used. That can help developers find a costly loop or a failing tool call without mistaking fewer calls for a better agent.

The practical payoff is sharper debugging and more meaningful comparisons between changes to an agent’s harness. It is not a magic quality score: results still depend on the task, verifier and test conditions, and the report’s examples do not establish that every agent will improve.

Our read

This is a useful bit of plumbing with a visible purpose. If your agent reaches the right answer by repeatedly fetching the same file, “success” is technically true and operationally rather cheeky. Inspecting the run makes that kind of waste easier to spot, while the verifier helps keep optimisation from becoming a race to do less work regardless of the result.

What to watch

  • Whether Hermes and NeMo Relay’s tracing approach is usable beyond the documented example.
  • Whether future ToolPerf results publish enough detail for others to reproduce the comparisons.
  • Whether trace-based evaluation leads to better verified outcomes, not merely shorter runs.

Discussion spark: When judging an AI agent, should a correct final answer count as success if the trace shows wasteful retries or hidden failures?

Sources and evidence

Nous/Hermes Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by Nous Research.