yesod.work

Model economics / detail

deepseek-v4-proA

Dispatch tier: Inexpensive (complexity class this model is routed for)

The fleet's workhorse: 845 runs — triple anyone else — at 0.71, matching the best high-volume peers while eating 771M tokens without flinching. Not flawless, just relentlessly dependable at scale.

The measured record

What the ledger shows

Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.

Agent runs

845

Success rate

71%

Metered spend

$198.83

674 of 845 runs metered

Cost / metered run

$0.30

Tokens in

770.9M

across 674 token-measured runs

Tokens out

8.7M

Tokens in / run

1.1M

Tokens out / run

13.0k

Outcomes

Status breakdown

Every run ends in exactly one status. The bar is the whole record, to scale.

success · 598 (70.8%)failure · 201 (23.8%)killed · 7 (0.8%)timeout · 10 (1.2%)lost · 29 (3.4%)

Run duration

How long the runs take

Wall-clock distribution across measured runs, from the fastest exit to the longest grind.

Min

1s

p25

39s

Median

4m

p75

9m

p90

16m

p99

71m

Max

2.2h

Mean

8m

Duration measured on 845 of 845 runs.

Commentary

Idiosyncrasies

What stands out in this model's numbers — shape, appetite, and failure habits.

Commentary

Lessons learned

Practical routing and operations takeaways, grounded in the same record.