HarnessRouter Benchmark 01

Same task. Eight configurations. About a 475× cost range.

HarnessRouter recorded the same controlled synthetic-data Care Prep task and input across Codex, Claude Code, and Hermes configurations. Measured cost ranged from 0.47 to 223 credits, while end-to-end latency ranged from 1m 25s to 4m 36s.

This is one controlled test, not a universal ranking. It is the first in a series of HarnessRouter benchmarks. We will publish more tasks, configurations, and evaluation methods over time.

Configurations
8

Harness × model combinations

Measured cost range
≈475×

0.47 to 223 credits

Latency range
3.2×

1m 25s to 4m 36s

Eight configurations, one task

Results are grouped by model so the effect of the harness remains visible. Lower cost and lower latency are better within this test.

Care Prep benchmark results for eight harness and model configurations
HarnessModelEnd-to-end latencyCost
Hermesgpt-5.22m 33s0.47 credits
Codexgpt-5.24m 36s0.72 credits
Hermesgpt-5.51m 25s40.7 credits
Codexgpt-5.52m 46s59.7 credits
Hermesclaude-sonnet-4.62m 49s39.6 credits
Claude Codeclaude-sonnet-4.61m 31s83.5 credits
Hermesclaude-opus-4.82m 42s150 credits
Claude Codeclaude-opus-4.83m 18s223 credits

The approximately 475× figure compares 223 with 0.47 credits. Because the displayed measurements are rounded, the ratio is also approximate. The latency range is 276 ÷ 85, or approximately 3.2×. The lowest-cost configuration consumed 99.8% fewer credits than the highest-cost configuration.

The choice changes with the objective

A single popularity rank cannot answer which setup is right for a production workload. Cost and speed can move in different directions.

Lowest measured costHermes + gpt-5.2

0.47 credits, compared with 223 credits for the highest-cost configuration in this test.

Lowest measured latencyHermes + gpt-5.5

1m 25s end to end, compared with 4m 36s for the slowest configuration in this test.

A visible tradeoffClaude Sonnet 4.6

Claude Code finished faster, while Hermes consumed fewer credits. The right choice depends on the workload objective.

The source run capture

The result table above transcribes the eight configurations shown in this HarnessRouter capture.

HarnessRouter Care Prep benchmark showing eight Codex, Claude Code, and Hermes configurations with recorded latency and cost
HarnessRouter Care Prep benchmark capture, August 2026. Harness identifiers are truncated in the product interface. Open the full-size capture.

The same experiment also tested grounding.

Across the same eight configurations, the experiment retained 5 observed runs per configuration while keeping the dataset, task, Skill, schema, and output contract fixed.

HarnessRouter controlled synthetic-data evaluation showing execution success, strict grounding pass rates, and p95 duration for eight harness and model configurations
Observed history from the same controlled synthetic-data experiment. “Strict grounding pass” checks schema compliance, exact supplied facts, evidence-reference validity, and required human-review flag. It is not a composite quality score or clinical validation.

How to read these results

This publication is designed to make the measurement and its limits clear, not to declare one harness universally best.

Test scope
One controlled synthetic-data Care Prep experiment across eight harness and model configurations.
Observed history
The experiment retained 5 observed runs per configuration and applied an objective validator to completed outputs.
Held constant
The dataset, Task, Skill, schema, and output contract remained fixed across configurations.
Changed
The selected agent harness and model configuration changed between rows.
Measurements
The source capture records cost and end-to-end latency for each configuration. The observed history reports execution success, strict grounding pass rate, and p95 duration.
Quality boundary
This publication does not combine quality into one score. The objective validator reports specific grounding checks; safety and workflow usefulness remain separate human-review criteria.
Interpretation
These results apply to this task and configuration set. Results can change with the task, input, harness, model, environment, and configuration.

Compare harnesses on the work your product actually runs.

HarnessRouter Cloud records execution traces across harness and model configurations so teams can compare cost, latency, and quality using their own production criteria.