yesod.work

Model economics

Every model.
Every run on record.

This is Stephen's personal Yesod instance — he has been using it to build Yesod itself and his other tools, and this scoreboard is what he has found so far: what each model actually did here — runs, outcomes, spend — with a verdict grounded in that record, not in marketing.

The historical record

Every model the factory has run

Autonomous agent runs only — excludes process pours (ys-proc-) and operator activity. Durations use COALESCE(finished_at, now) for runs still in progress.Token/cost aggregates exclude runs with no captured token data (cost-blind hosts leave tokens NULL).

ModelTierRunsStatus breakdownSuccessTokens inTokens outMedian runAI verdict
deepseek-v4-proA84571%770.9M8.7M4mThe fleet's workhorse: 845 runs — triple anyone else — at 0.71, matching the best high-volume peers while eating 771M tokens without flinching. Not flawless, just relentlessly dependable at scale.
glm-5p2A36667%773.2M5.4M6mSecond-biggest workhorse in the fleet, and once you strip the 55 infra-lost runs its 79% adjusted success rate quietly leads the high-volume pack. Docked from S for gulping ~2M tokens per job — reliable, but thirsty.
kimi-k2p7-codeA29372%420.6M2.6M4m293 runs at 72% (76% excluding infra-lost), 4m median, lean token diet — quietly outperforms every other high-volume workhorse. Not flashy enough for S, but the fleet's best price-per-success.
claude-opus-4-8B27465%660.7M4.9M6mReal workhorse volume (274 runs) and ~72% success once you forgive the 27 infra-lost runs — but it burns 2.4M tokens per job to match what kimi-k2p7 does on half the budget. Solid, not thrifty.
gpt-5.6-solC13643%5m136 runs is a real sample, and 0.43 is a real problem — worst rate of any high-volume model. The merge-conflict beat is genuinely brutal, which buys it a C instead of a D, but hazard pay isn't a grade.
claude-fable-5A5671%61.5M399.6k18mMatches the fleet leaders' ~71% (75% excluding lost runs) while grinding 18-minute jobs — triple the median workload. 56 runs is a real record, not an anecdote. The heavy-lifter that doesn't drop the bar.
kimi-k2p6S54100%59.3M435.1k5m54 runs, 54 successes, zero kills, timeouts, or losses — a flawless sheet at real volume, done cheap and fast. Others hauled more freight; nobody else hauled it without dropping a single crate.
claudeC4151%58.2M436.6k37s41 runs, 51% success, 37-second median — a coin flip that quits before its coffee cools. Enough volume to trust the number, and the number says: route this alias to a real model.
kimi-k2.6C2556%3m25 runs at 56% (61% excusing infra losses) — below the fleet's ~0.7 workhorse line with a sample too thin to blame bad luck. Its sibling k2p6 went 54-for-54; this one is the family's rough draft.
kimi-k3B1771%17.4M93.5k47mMatches the fleet-leading 0.71 success rate on 17 slow, 47-minute grinds, and it's the merge-conflict fixer of last resort. Reliable specialist; needs more reps before the tortoise gets a medal.
gpt-5.6-lunaA12100%6mTwelve for twelve, zero drama — a flawless record that just clears the anecdote threshold. Perfection at boutique volume earns an A; ship it a few hundred runs and we'll talk S.
kimi-k2p6-turboC1267%29sTwelve runs at 0.67 with a 29-second median: it sprints where its sibling k2p6 strolls to a perfect 54/54. Speed is charming, but a fast coin-flip-plus is still thin evidence. Come back with volume.
codexB8100%2mEight for eight is a spotless sheet, but eight runs is a weekend hobby next to deepseek's 845. Perfect record, anecdotal volume — B, with an open invitation to show up more often.
minimax-m3C475%9.0M47.0k5mThree wins in four at-bats is a nice afternoon, not a career. The 0.75 rate would look respectable at fleet scale, but four runs is an anecdote — come back with a real sample.
claude-4.8-opusD333%25sThree misrouted dispatches under a name that doesn't exist, one lucky success, 25s median. This isn't a model record, it's a typo with a scoreboard entry. Fix the router; judge the real claude-opus-4-8 instead.
gpt-5.6-terraC367%14mTwo-for-two on runs the infra didn't eat, which is technically a perfect record and statistically a coin flip. Three runs is a cameo, not a career — come back with volume.
deepseek-v3D20%008sTwo runs, zero tokens, dead in 8 seconds each — this isn't a model, it's a misconfigured endpoint wearing a name tag. 0% success is unforgiving, even if the wiring, not the weights, deserves the blame.
qwen3p6-plusC2100%947.9k9.3k2mTwo runs, two wins, barely a million tokens burned — a flawless cameo, not a career. Perfect record, anecdotal volume. Come back after 50 runs and we'll talk tiers.
claude-opusC10%408.0hOne run, lost by the infra, then left 'running' for 17 days — a 24,482-minute monument to nobody checking on it. Zero evidence of skill or failure; C by presumption of innocence, not merit.
claude-opus-5C10%9sOne run, dead in nine seconds — a dispatch typo cosplaying as a model. No evidence of skill or its absence; C as a shrug, not a sentence. Fix the label and try again.

Fleet-wide aggregates from the factory's run ledger — model names and counts only; no notes, projects, or hosts. Token columns count only runs with captured token data. Tiers and verdicts are AI-generated and unapologetically subjective: one judge agent per model, each shown the full fleet record and asked to grade its model relative to peers, with small samples capped. The measured columns are the facts; the verdict is commentary.