Model economics / detail
kimi-k3B
Dispatch tier: Medium (complexity class this model is routed for)
Matches the fleet-leading 0.71 success rate on 17 slow, 47-minute grinds, and it's the merge-conflict fixer of last resort. Reliable specialist; needs more reps before the tortoise gets a medal.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
17
Success rate
71%
Metered spend
Not observed
cost-blind: no run was metered
Cost / metered run
Not observed
Tokens in
17.4M
across 2 token-measured runs
Tokens out
93.5k
Tokens in / run
8.7M
Tokens out / run
46.7k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 4 of 17 runs (23.5% of total)
- timeout: 1 of 17 runs (5.9% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
11s
p25
34m
Median
47m
p75
77m
p90
1.9h
p99
2.7h
Max
2.9h
Mean
56m
Duration measured on 17 of 17 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- Slowest median in the fleet by far: 47 minutes typical, with p25 already past half an hour — these are all long hauls.
- Token telemetry on just 2 of 17 runs, and those two averaged 8.7M tokens in — the biggest measured appetite anywhere, on the thinnest coverage.
- 12-of-17 (71%) matches the fleet-leading rate despite the workload.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Fleet-best success on the longest grinds makes it the heavy-job specialist its medium-tier routing implies.
- With 2 metered runs in 17, its economics are a rumor — instrument before scaling it up.