yesod.work

Model economics / detail

claude-opus-4-8B

Dispatch tier: Expensive (complexity class this model is routed for)

Real workhorse volume (274 runs) and ~72% success once you forgive the 27 infra-lost runs — but it burns 2.4M tokens per job to match what kimi-k2p7 does on half the budget. Solid, not thrifty.

The measured record

What the ledger shows

Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.

Agent runs

274

Success rate

65%

Metered spend

$584.11

172 of 274 runs metered

Cost / metered run

$3.40

Tokens in

660.7M

across 172 token-measured runs

Tokens out

4.9M

Tokens in / run

3.8M

Tokens out / run

28.4k

Outcomes

Status breakdown

Every run ends in exactly one status. The bar is the whole record, to scale.

success · 178 (65.0%)failure · 64 (23.4%)killed · 1 (0.4%)timeout · 4 (1.5%)lost · 27 (9.9%)

Run duration

How long the runs take

Wall-clock distribution across measured runs, from the fastest exit to the longest grind.

Min

4s

p25

28s

Median

6m

p75

11m

p90

19m

p99

49m

Max

60m

Mean

8m

Duration measured on 274 of 274 runs.

Commentary

Idiosyncrasies

What stands out in this model's numbers — shape, appetite, and failure habits.

Commentary

Lessons learned

Practical routing and operations takeaways, grounded in the same record.