Model economics / detail
kimi-k2p7-codeA
Dispatch tier: Inexpensive (complexity class this model is routed for)
293 runs at 72% (76% excluding infra-lost), 4m median, lean token diet — quietly outperforms every other high-volume workhorse. Not flashy enough for S, but the fleet's best price-per-success.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
293
Success rate
72%
Metered spend
$76.72
157 of 293 runs metered
Cost / metered run
$0.49
Tokens in
420.6M
across 191 token-measured runs
Tokens out
2.6M
Tokens in / run
2.2M
Tokens out / run
13.6k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 60 of 293 runs (20.5% of total)
- timeout: 5 of 293 runs (1.7% of total)
- lost: 17 of 293 runs (5.8% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
2s
p25
53s
Median
4m
p75
11m
p90
24m
p99
60m
Max
60m
Mean
9m
Duration measured on 293 of 293 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- Best success rate (72%) of any model with 250+ runs, on an inexpensive-tier ticket of ~$0.49 per metered run.
- The duration tail stops dead: p99 and max are the same 3,611 seconds — nothing ever ran past the hour mark.
- Half of claude-opus-4-8’s output tokens per run at roughly a seventh of the cost per run, with a slightly better success rate.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Pound-for-pound the strongest workhorse in the fleet — the inexpensive routing tier is flattering it, not constraining it.
- 17 infra-lost runs (76% success excluding them) — most of its remaining headroom is operational, not model quality.