Model economics
Every model.
Every run on record.
This is Stephen's personal Yesod instance — he has been using it to build Yesod itself and his other tools, and this scoreboard is what he has found so far: what each model actually did here — runs, outcomes, spend — with a verdict grounded in that record, not in marketing.
The historical record
Every model the factory has run
Autonomous agent runs only — excludes process pours (ys-proc-) and operator activity. Durations use COALESCE(finished_at, now) for runs still in progress.Token/cost aggregates exclude runs with no captured token data (cost-blind hosts leave tokens NULL).
| Model | Tier | Runs | Status breakdown | Success | Tokens in | Tokens out | Median run | AI verdict |
|---|---|---|---|---|---|---|---|---|
| deepseek-v4-pro | A | 845 | 71% | 770.9M | 8.7M | 4m | The fleet's workhorse: 845 runs — triple anyone else — at 0.71, matching the best high-volume peers while eating 771M tokens without flinching. Not flawless, just relentlessly dependable at scale. | |
| glm-5p2 | A | 366 | 67% | 773.2M | 5.4M | 6m | Second-biggest workhorse in the fleet, and once you strip the 55 infra-lost runs its 79% adjusted success rate quietly leads the high-volume pack. Docked from S for gulping ~2M tokens per job — reliable, but thirsty. | |
| kimi-k2p7-code | A | 293 | 72% | 420.6M | 2.6M | 4m | 293 runs at 72% (76% excluding infra-lost), 4m median, lean token diet — quietly outperforms every other high-volume workhorse. Not flashy enough for S, but the fleet's best price-per-success. | |
| claude-opus-4-8 | B | 274 | 65% | 660.7M | 4.9M | 6m | Real workhorse volume (274 runs) and ~72% success once you forgive the 27 infra-lost runs — but it burns 2.4M tokens per job to match what kimi-k2p7 does on half the budget. Solid, not thrifty. | |
| gpt-5.6-sol | C | 136 | 43% | — | — | 5m | 136 runs is a real sample, and 0.43 is a real problem — worst rate of any high-volume model. The merge-conflict beat is genuinely brutal, which buys it a C instead of a D, but hazard pay isn't a grade. | |
| claude-fable-5 | A | 56 | 71% | 61.5M | 399.6k | 18m | Matches the fleet leaders' ~71% (75% excluding lost runs) while grinding 18-minute jobs — triple the median workload. 56 runs is a real record, not an anecdote. The heavy-lifter that doesn't drop the bar. | |
| kimi-k2p6 | S | 54 | 100% | 59.3M | 435.1k | 5m | 54 runs, 54 successes, zero kills, timeouts, or losses — a flawless sheet at real volume, done cheap and fast. Others hauled more freight; nobody else hauled it without dropping a single crate. | |
| claude | C | 41 | 51% | 58.2M | 436.6k | 37s | 41 runs, 51% success, 37-second median — a coin flip that quits before its coffee cools. Enough volume to trust the number, and the number says: route this alias to a real model. | |
| kimi-k2.6 | C | 25 | 56% | — | — | 3m | 25 runs at 56% (61% excusing infra losses) — below the fleet's ~0.7 workhorse line with a sample too thin to blame bad luck. Its sibling k2p6 went 54-for-54; this one is the family's rough draft. | |
| kimi-k3 | B | 17 | 71% | 17.4M | 93.5k | 47m | Matches the fleet-leading 0.71 success rate on 17 slow, 47-minute grinds, and it's the merge-conflict fixer of last resort. Reliable specialist; needs more reps before the tortoise gets a medal. | |
| gpt-5.6-luna | A | 12 | 100% | — | — | 6m | Twelve for twelve, zero drama — a flawless record that just clears the anecdote threshold. Perfection at boutique volume earns an A; ship it a few hundred runs and we'll talk S. | |
| kimi-k2p6-turbo | C | 12 | 67% | — | — | 29s | Twelve runs at 0.67 with a 29-second median: it sprints where its sibling k2p6 strolls to a perfect 54/54. Speed is charming, but a fast coin-flip-plus is still thin evidence. Come back with volume. | |
| codex | B | 8 | 100% | — | — | 2m | Eight for eight is a spotless sheet, but eight runs is a weekend hobby next to deepseek's 845. Perfect record, anecdotal volume — B, with an open invitation to show up more often. | |
| minimax-m3 | C | 4 | 75% | 9.0M | 47.0k | 5m | Three wins in four at-bats is a nice afternoon, not a career. The 0.75 rate would look respectable at fleet scale, but four runs is an anecdote — come back with a real sample. | |
| claude-4.8-opus | D | 3 | 33% | — | — | 25s | Three misrouted dispatches under a name that doesn't exist, one lucky success, 25s median. This isn't a model record, it's a typo with a scoreboard entry. Fix the router; judge the real claude-opus-4-8 instead. | |
| gpt-5.6-terra | C | 3 | 67% | — | — | 14m | Two-for-two on runs the infra didn't eat, which is technically a perfect record and statistically a coin flip. Three runs is a cameo, not a career — come back with volume. | |
| deepseek-v3 | D | 2 | 0% | 0 | 0 | 8s | Two runs, zero tokens, dead in 8 seconds each — this isn't a model, it's a misconfigured endpoint wearing a name tag. 0% success is unforgiving, even if the wiring, not the weights, deserves the blame. | |
| qwen3p6-plus | C | 2 | 100% | 947.9k | 9.3k | 2m | Two runs, two wins, barely a million tokens burned — a flawless cameo, not a career. Perfect record, anecdotal volume. Come back after 50 runs and we'll talk tiers. | |
| claude-opus | C | 1 | 0% | — | — | 408.0h | One run, lost by the infra, then left 'running' for 17 days — a 24,482-minute monument to nobody checking on it. Zero evidence of skill or failure; C by presumption of innocence, not merit. | |
| claude-opus-5 | C | 1 | 0% | — | — | 9s | One run, dead in nine seconds — a dispatch typo cosplaying as a model. No evidence of skill or its absence; C as a shrug, not a sentence. Fix the label and try again. |
Fleet-wide aggregates from the factory's run ledger — model names and counts only; no notes, projects, or hosts. Token columns count only runs with captured token data. Tiers and verdicts are AI-generated and unapologetically subjective: one judge agent per model, each shown the full fleet record and asked to grade its model relative to peers, with small samples capped. The measured columns are the facts; the verdict is commentary.