The announcement
Harness Arena, for production tasks
HarnessRouter today introduces Harness Arena: run the same production task across agent harness and model configurations, score the results on success and quality first and cost and latency second, and send your traffic to the winner.
The idea is simple and overdue. Agent harnesses such as Codex, Claude Code, and Hermes are complete, continuously improving systems, and on HarnessRouter each harness runs with multiple models. That makes the real unit of choice a configuration, a harness together with a model at minimum, and there are more viable configurations than teams typically evaluate by hand.
Harness Arena makes that choice competitive instead of assumed. Every configuration runs the same task under the same contract, so the results are comparable, and the comparison is about your task.
The problem
The default configuration is quietly expensive
Teams commonly choose a harness once, early, and by familiarity. The choice then hardens into architecture and is rarely revisited, even as harnesses and models keep shipping updates.
The spread that decision leaves on the table is measurable. In HarnessRouter's published same-task benchmark, eight harness and model configurations each ran the same task five times on identical input: cost per task varied by approximately 475 times, and p95 end-to-end latency varied by more than 3 times. Results vary by task, and that is exactly the point: the way to know your task's spread is to run your task.
A default configuration is not neutral. On some task classes it is the best available choice; on others it silently taxes every single run. Without competition, it is hard to tell which situation you are in.
The mechanism
How the arena works
Harness Arena applies the same method behind HarnessRouter's published benchmarks to your own workload:
- Bring a real task: a production fixture with real input, not a synthetic puzzle.
- Field the contenders: the harness and model configurations you want to compare, run through one API.
- Gate on success first: task success and required quality are constraints, not weights. A cheap failure is not a bargain.
- Compare cost and latency among the configurations that pass the gate.
- Let the winner take the traffic: point the task class at the winning configuration, and rerun the arena when harnesses or models update.
The unit that comes out of an arena run is the one that matters in production: cost per successful task, at a latency you can live with.
The foundation
Competition needs a level field
A comparison is only as trustworthy as its control variables. Harness Arena is possible because every contender runs behind the same open contract: the Unified Harness Protocol (UHP), specified at unifiedharnessprotocol.org, covers tasks, sessions, streaming, files, and artifacts identically for every harness. Same task, same input, same contract; the variable under test is the configuration.
This is also what separates an arena for harnesses from a model leaderboard. Public leaderboards rank models in isolation, and they are useful for that. A production agent task runs a complete harness: the loop, the tools, the file handling, the recovery. Ranking models and racing configurations are different measurements, and production teams need the second.
What changes
Harness choice becomes a standing decision
With an arena in the loop, harness selection stops being a one-time bet and becomes an operating practice: per task class, evidence-based, and rerunnable whenever the field changes. New model release? Rerun the arena. New harness on the platform? It is available to field in the next race.
Serving the winner does not require new plumbing. Your app sends one task; HarnessRouter runs the best harness in a sandbox and returns renderable artifacts to your UI. Your arena results tell you which configuration that is for each class of work; point the task class at it, and the same integration carries the decision.
For the broader field, HarnessRouter's public rankings track how harnesses perform against each other over time. The arena is the same discipline turned inward: the public rankings show the field, and your arena runs deliver the verdict for your product.
The conclusion
Every production task deserves a competition
Agent harnesses keep getting better, at different rates, in different directions. The strategy that ages well in that landscape is not picking the right harness once; it is keeping the competition open so the right harness keeps winning your tasks.
That is what Harness Arena is for. Bring one real task and let the field race for it.
FAQ
Harness Arena FAQ
What is Harness Arena?
Harness Arena is HarnessRouter's way of making harness choice competitive. You field several agent harness and model configurations against one real task under identical conditions; anything that misses your success bar is out, the rest are ranked on cost and latency, and the winner earns that task class.
How is Harness Arena different from a model leaderboard?
A model leaderboard scores models by themselves, usually on public prompts. An arena run scores complete configurations, harness plus model, doing your actual work, in the unit production cares about: cost per successful task. One tells you which model is strong; the other tells you what to ship.
Can I run Harness Arena on my own tasks?
Yes, that is the point. Any real task from your product can become an arena: identical input to every configuration, your own success criteria as the gate. The public HarnessRouter benchmarks follow this method too, so your results sit directly beside the published evidence.
Put your harnesses in the arena
Create an API key, field two or three harness and model configurations on one real task, and route your traffic to the configuration that earns it.
Start building with the winner

