yesod.work

Model economics / detail

gpt-5.6-lunaA

Dispatch tier: Inexpensive (complexity class this model is routed for)

Twelve for twelve, zero drama — a flawless record that just clears the anecdote threshold. Perfection at boutique volume earns an A; ship it a few hundred runs and we'll talk S.

The measured record

What the ledger shows

Fleet-ledger aggregates for this model only — run counts, outcomes, spend, and token appetite. Missing telemetry reads “Not observed”, never zero.

Agent runs

12

Success rate

100%

Metered spend

Not observed

cost-blind: no run was metered

Cost / metered run

Not observed

Tokens in

Not observed

Tokens out

Not observed

Tokens in / run

Not observed

Tokens out / run

Not observed

Outcomes

Status breakdown

Every run ends in exactly one status. The bar is the whole record, to scale.

success · 12 (100.0%)

No non-success statuses on record — every run succeeded.

Run duration

How long the runs take

Wall-clock distribution across measured runs, from the fastest exit to the longest grind.

Min

5m

p25

6m

Median

6m

p75

7m

p90

7m

p99

7m

Max

7m

Mean

6m

Duration measured on 12 of 12 runs.

Commentary

Idiosyncrasies

What stands out in this model's numbers — shape, appetite, and failure habits.

Commentary

Lessons learned

Practical routing and operations takeaways, grounded in the same record.