Nous Research Launches Hermes Index for Agent Cost and Performance
Nous Research has published the Hermes Index, a public leaderboard that compares models inside its Hermes Agent framework running at maximum reasoning depth. The score combines four evaluations: Terminal-Bench 4, Terminal-Bench Science, SkillsBench, and the new Hermes Bench, where an agent works through 150 varied tasks in a real workspace using files and external tools.
- Top of the board: Claude Opus 5.5 took first place at roughly $5 per task, ahead of GPT-6 Astra, which placed second and was the most expensive entry at $11.60 per task. Claude Sonnet 5.5 came third at $2.80.
- Best value: DeepSeek V4.1 Flash (36.91 points at $0.26 per task) and GPT-6 Luna ($0.14 per task) landed on the Pareto frontier for cost against productivity.
The index is set up as a cost-aware ranking rather than a raw capability chart, so the cheapest model is not automatically the best one for a given workload.
Related: Nous Research Unveils Hermes Desktop, OpenAI Releases GPT-6 Sol and GPT-6 Luna, ApprenticeBench Tests AI’s Ability to Learn and Work Like a New Employee