Model economics / detail
deepseek-v4-proA
Dispatch tier: Inexpensive (complexity class this model is routed for)
The fleet's workhorse: 845 runs — triple anyone else — at 0.71, matching the best high-volume peers while eating 771M tokens without flinching. Not flawless, just relentlessly dependable at scale.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
845
Success rate
71%
Metered spend
$198.83
674 of 845 runs metered
Cost / metered run
$0.30
Tokens in
770.9M
across 674 token-measured runs
Tokens out
8.7M
Tokens in / run
1.1M
Tokens out / run
13.0k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 201 of 845 runs (23.8% of total)
- killed: 7 of 845 runs (0.8% of total)
- timeout: 10 of 845 runs (1.2% of total)
- lost: 29 of 845 runs (3.4% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
1s
p25
39s
Median
4m
p75
9m
p90
16m
p99
71m
Max
2.2h
Mean
8m
Duration measured on 845 of 845 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- The tail is the story: median run is ~4 minutes, p99 is ~71 minutes — a 17x spread, with one run wandering out to 2.2 hours.
- Leanest token diet of the big workhorses: ~1.1M tokens in per run, versus 2.9M for glm-5p2 and 3.8M for claude-opus-4-8.
- 201 outright failures is the largest raw failure pile in the fleet — but at 845 runs the 71% rate is the fleet-leader number, not a fluke.
- Cheapest metered workhorse at roughly $0.30 per costed run.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- The inexpensive-tier default routing is earned: fleet-leading success at fleet-leading volume and the lowest cost per run.
- Budget wall-clock for the tail — a scheduler that assumes the 4-minute median will be surprised by the 71-minute p99.
- At $0.30 a run, retrying its failures costs less than one claude-opus-4-8 attempt.