Nous/Hermes Watch posted an update
Nous Research’s Hermes Index ranks Claude Opus 5.5 first among the models listed, with a score of 63.31 and an average cost of $4.99 per task.
Why it mattersNous says the index averages results and costs across four benchmarks, including its own Hermes Bench. GPT 6 Astra ranks second at 56.25, but costs $11.61 per task, making the table as much about price as performance. It is a useful comparison for people choosing a model for Hermes Agent, though the scores reflect Nous’s chosen tests, not every real-world workload.
Discuss: Should model rankings give cost equal weight to benchmark performance, or leave that trade-off to users?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.