Model economics / detail
minimax-m3C
Dispatch tier: Inexpensive (complexity class this model is routed for)
Three wins in four at-bats is a nice afternoon, not a career. The 0.75 rate would look respectable at fleet scale, but four runs is an anecdote — come back with a real sample.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
4
Success rate
75%
Metered spend
Not observed
cost-blind: no run was metered
Cost / metered run
Not observed
Tokens in
9.0M
across 4 token-measured runs
Tokens out
47.0k
Tokens in / run
2.2M
Tokens out / run
11.8k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 1 of 4 runs (25.0% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
24s
p25
4m
Median
5m
p75
6m
p90
7m
p99
8m
Max
8m
Mean
5m
Duration measured on 4 of 4 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- 3-of-4 with full token coverage (~2.2M in per run) but zero cost telemetry.
- Durations scattered from 24 seconds to 8 minutes — no shape to speak of at n=4.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Four runs teach nothing about the model; leave it in the inexpensive rotation and let the sample accumulate.
- It already reports tokens — wiring up cost capture would make its first real sample actually priceable.