Model economics / detail
kimi-k2.6C
Dispatch tier: Inexpensive (complexity class this model is routed for)
25 runs at 56% (61% excusing infra losses) — below the fleet's ~0.7 workhorse line with a sample too thin to blame bad luck. Its sibling k2p6 went 54-for-54; this one is the family's rough draft.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
25
Success rate
56%
Metered spend
Not observed
cost-blind: no run was metered
Cost / metered run
Not observed
Tokens in
Not observed
Tokens out
Not observed
Tokens in / run
Not observed
Tokens out / run
Not observed
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 9 of 25 runs (36.0% of total)
- lost: 2 of 25 runs (8.0% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
3s
p25
8s
Median
3m
p75
5m
p90
8m
p99
12m
Max
12m
Mean
3m
Duration measured on 25 of 25 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- A quarter of its runs die almost instantly — p25 is 8 seconds against a 3.4-minute median.
- 14-of-25 (56%, 61% excusing 2 infra losses) while sibling k2p6 sits at a perfect 54-for-54.
- No token or cost telemetry on any of its 25 runs.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- 25 runs is thin, but trailing the fleet’s ~0.7 line while a near-identical sibling runs clean points at dispatch config, not family genetics.
- Chase the 8-second deaths first — if those are misfires, the real rate is closer to respectable.