Model economics / detail
gpt-5.6-lunaA
Dispatch tier: Inexpensive (complexity class this model is routed for)
Twelve for twelve, zero drama — a flawless record that just clears the anecdote threshold. Perfection at boutique volume earns an A; ship it a few hundred runs and we'll talk S.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
12
Success rate
100%
Metered spend
Not observed
cost-blind: no run was metered
Cost / metered run
Not observed
Tokens in
Not observed
Tokens out
Not observed
Tokens in / run
Not observed
Tokens out / run
Not observed
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
No non-success statuses on record — every run succeeded.
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
5m
p25
6m
Median
6m
p75
7m
p90
7m
p99
7m
Max
7m
Mean
6m
Duration measured on 12 of 12 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- 12-for-12 inside the tightest duration band in the fleet: every run between 281 and 445 seconds. A metronome.
- Zero token or cost telemetry across all 12 runs.
- 12 runs clears the anecdote bar, barely — one bad week would move the rate 8 points.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Perfect and predictable is exactly what you want from the inexpensive tier — feed it volume and see if the metronome keeps time.
- It is completely cost-blind, so "inexpensive" is currently a routing label, not a measurement.