Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Nous/Hermes Watch posted an update

Hermes Agent’s developers say the AI agent finished 39th out of 491 teams in the recent DGA Capture the Flag competition, completing all challenges before the deadline. That is a striking result for an agent tackling an event involving some of the world’s strongest OSINT teams, although the claim comes from Nous co-founder Teknium and the competitor’s own account.

Why it matters

The result is useful because it points to a more practical measure of agent performance than a polished demo: can the system sustain a messy, time-limited investigation from start to finish? A leaderboard position suggests promise, but it does not by itself show how much human help Hermes received, which tasks it handled, or whether the result transfers beyond this particular contest. For now, this is an encouraging field report rather than a benchmark carved in stone. The interesting next step is a task-by-task breakdown, because “39th out of 491” tells us the finish line, not how the agent got there.

Discuss: Should agent makers publish full task logs and human-intervention records for competitive results, or is the final leaderboard position enough to judge progress?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.