yesod.work

Model economics / detail

claude-4.8-opusD

Dispatch tier: Expensive (complexity class this model is routed for)

Three misrouted dispatches under a name that doesn't exist, one lucky success, 25s median. This isn't a model record, it's a typo with a scoreboard entry. Fix the router; judge the real claude-opus-4-8 instead.

The measured record

What the ledger shows

Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.

Agent runs

3

Success rate

33%

Metered spend

Not observed

cost-blind: no run was metered

Cost / metered run

Not observed

Tokens in

Not observed

Tokens out

Not observed

Tokens in / run

Not observed

Tokens out / run

Not observed

Outcomes

Status breakdown

Every run ends in exactly one status. The bar is the whole record, to scale.

success · 1 (33.3%)failure · 2 (66.7%)

Run duration

How long the runs take

Wall-clock distribution across measured runs, from the fastest exit to the longest grind.

Min

19s

p25

22s

Median

25s

p75

39s

p90

48s

p99

53s

Max

53s

Mean

32s

Duration measured on 3 of 3 runs.

Commentary

Idiosyncrasies

What stands out in this model's numbers — shape, appetite, and failure habits.

Commentary

Lessons learned

Practical routing and operations takeaways, grounded in the same record.