Model economics / detail
claude-fable-5A
Dispatch tier: Expensive (complexity class this model is routed for)
Matches the fleet leaders' ~71% (75% excluding lost runs) while grinding 18-minute jobs — triple the median workload. 56 runs is a real record, not an anecdote. The heavy-lifter that doesn't drop the bar.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
56
Success rate
71%
Metered spend
$105.84
9 of 56 runs metered
Cost / metered run
$11.76
Tokens in
61.5M
across 9 token-measured runs
Tokens out
399.6k
Tokens in / run
6.8M
Tokens out / run
44.4k
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
- failure: 10 of 56 runs (17.9% of total)
- killed: 1 of 56 runs (1.8% of total)
- timeout: 2 of 56 runs (3.6% of total)
- lost: 3 of 56 runs (5.4% of total)
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
3s
p25
8m
Median
18m
p75
27m
p90
36m
p99
44m
Max
45m
Mean
18m
Duration measured on 56 of 56 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- Longest jobs of any real-volume model: 18-minute median, p25 already at 7.6 minutes — nothing about it is quick.
- Telemetry covers only 9 of 56 runs, and those metered runs averaged $11.76 — the priciest per-run figure in the fleet where measured.
- Holds the fleet-leader ~71% success rate (75% excluding 3 infra-lost) while carrying triple-length workloads.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- The right destination for the long grind: it keeps fleet-best success on jobs other models never even attempt at that length.
- Do not trust its cost math yet — 9 metered runs out of 56 is a keyhole, and the view through it is expensive.