Model economics / detail
claudeC
Dispatch tier: Unclassified (complexity class this model is routed for)
41 runs, 51% success, 37-second median — a coin flip that quits before its coffee cools. Enough volume to trust the number, and the number says: route this alias to a real model.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
41
Success rate
51%
Metered spend
$49.34
21 of 41 runs metered
Cost / metered run
$2.35
Tokens in
58.2M
across 21 token-measured runs
Tokens out
436.6k
Tokens in / run
2.8M
Tokens out / run
20.8k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 18 of 41 runs (43.9% of total)
- lost: 2 of 41 runs (4.9% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
5s
p25
17s
Median
37s
p75
9m
p90
13m
p99
20m
Max
21m
Mean
5m
Duration measured on 41 of 41 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- A generic harness alias, not a model — the unclassified tier is the router shrugging, and this record blends whatever it resolved to.
- 37-second median against a 9-minute p75: half its runs barely start before ending.
- $2.35 per metered run for a 51% success rate — coin-flip outcomes at mid-expensive prices.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Dispatch a specific model name; an alias’s scoreboard entry is unattributable by construction.
- The short median plus the 51% rate reads as early quits rather than hard-fought failures — a harness question before a model question.