yesod.work

Model economics / detail

kimi-k2.6C

Dispatch tier: Inexpensive (complexity class this model is routed for)

25 runs at 56% (61% excusing infra losses) — below the fleet's ~0.7 workhorse line with a sample too thin to blame bad luck. Its sibling k2p6 went 54-for-54; this one is the family's rough draft.

The measured record

What the ledger shows

Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.

Agent runs

25

Success rate

56%

Metered spend

Not observed

cost-blind: no run was metered

Cost / metered run

Not observed

Tokens in

Not observed

Tokens out

Not observed

Tokens in / run

Not observed

Tokens out / run

Not observed

Outcomes

Status breakdown

Every run ends in exactly one status. The bar is the whole record, to scale.

success · 14 (56.0%)failure · 9 (36.0%)lost · 2 (8.0%)

Run duration

How long the runs take

Wall-clock distribution across measured runs, from the fastest exit to the longest grind.

Min

3s

p25

8s

Median

3m

p75

5m

p90

8m

p99

12m

Max

12m

Mean

3m

Duration measured on 25 of 25 runs.

Commentary

Idiosyncrasies

What stands out in this model's numbers — shape, appetite, and failure habits.

Commentary

Lessons learned

Practical routing and operations takeaways, grounded in the same record.