Model economics / detail
codexB
Dispatch tier: Unclassified (complexity class this model is routed for)
Eight for eight is a spotless sheet, but eight runs is a weekend hobby next to deepseek's 845. Perfect record, anecdotal volume — B, with an open invitation to show up more often.
The measured record
What the ledger shows
Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.
Agent runs
8
Success rate
100%
Metered spend
Not observed
cost-blind: no run was metered
Cost / metered run
Not observed
Tokens in
Not observed
Tokens out
Not observed
Tokens in / run
Not observed
Tokens out / run
Not observed
Outcomes
Status breakdown
Every run ends in exactly one status. The bar is the whole record, to scale.
No non-success statuses on record — every run succeeded.
Run duration
How long the runs take
Wall-clock distribution across measured runs, from the fastest exit to the longest grind.
Min
89s
p25
2m
Median
2m
p75
3m
p90
6m
p99
7m
Max
7m
Mean
3m
Duration measured on 8 of 8 runs.
Commentary
Idiosyncrasies
What stands out in this model's numbers — shape, appetite, and failure habits.
- 8-for-8, every run inside a tidy 89-to-405-second window.
- A generic harness alias in the unclassified tier — the record cannot say which actual model earned the streak.
- No token or cost data captured.
Commentary
Lessons learned
Practical routing and operations takeaways, grounded in the same record.
- Spotless but anecdotal — eight runs is a weekend, not a track record.
- Name the underlying model in dispatch so future wins land on an attributable scoreboard row.