Model economics / detail
glm-5p2A
Dispatch tier: Medium (complexity class this model is routed for)
Second-biggest workhorse in the fleet, and once you strip the 55 infra-lost runs its 79% adjusted success rate quietly leads the high-volume pack. Docked from S for gulping ~2M tokens per job — reliable, but thirsty.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
366
Success rate
67%
Metered spend
$205.40
223 of 366 runs metered
Cost / metered run
$0.92
Tokens in
773.2M
across 268 token-measured runs
Tokens out
5.4M
Tokens in / run
2.9M
Tokens out / run
20.2k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 62 of 366 runs (16.9% of total)
- timeout: 2 of 366 runs (0.5% of total)
- lost: 55 of 366 runs (15.0% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
5s
p25
2m
Median
6m
p75
11m
p90
19m
p99
49m
Max
23.0h
Mean
12m
Duration measured on 366 of 366 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- 55 lost runs — 15% of its total and the worst infra-loss count in the fleet; the model was often not the thing that failed.
- One run logged 82,850 seconds (~23 hours), dwarfing its own p99 of ~49 minutes.
- Biggest input appetite of the volume models at ~2.9M tokens in per run, and about triple deepseek-v4-pro’s cost per costed run.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Excuse the infra losses and the adjusted rate (~79%) leads the high-volume pack — fix the harness before blaming the model.
- Put a runtime cap on its jobs; the 23-hour outlier says nothing else will.
- Its medium-tier routing is priced in: ~$0.92 per metered run buys workhorse volume, not a rate premium.