Which model is cheapest for what you actually ship?
Cost per solved task — what it takes to get a task finished, not the token price. Every number carries a source and a verified date.
I build tasks,
about a month.
multi-file refactor, feature with tests · Tier: moderate — measured rows unchanged.
DeepSeek V4 Flash costs $0.12 per solved task against Claude Opus 5 at $12.01 — ▼ 100x cheaper for 18 fewer points of pass rate. Over 200 tasks that is $2.4k a month.
Skip the ranked rowsMeasured · The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.
Modelled · Pass rate published; cost is Solvency's loop model at verified prices.
Stale pass rates (7) · Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.
Measured · The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.
Modelled · Pass rate published; cost is Solvency's loop model at verified prices.
Stale pass rates (7) · Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.
Not shown (8) · no published pass rate, reported as missing, never estimated
No published pass rate, so reported as missing rather than estimated: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.3 Codex, DeepSeek V4 Pro, Grok 4.6, Mistral Medium 3.5.
Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · Sources
Permalink with your name, re-priced when prices change.
Pro (soon): export CSV/JSON · price-change alerts · your own prices →
Assumptions (4) · move modelled rows only
These move modelled rows only. A model with no published cached-input price is computed uncached and says so in its row, rather than dropping out of its group. Measured rows carry a cost the benchmark observed and cannot be changed by any assumption. Frontier efficiency has no published source and is applied only to modelled rows — untick it to see unadjusted results. Methodology.
The frontier
Why not just pick the top score?
Lower and to the right is better. The line joins the measured models nobody beats on both axes.
Token price predicts almost nothing: DeepSeek V4 Flash is 19x cheaper per output token than Claude Opus 5 — and 100x cheaper per solved task.
Read the note →Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21
Modelled points move with your tier and assumptions; measured points never do. Hover or focus a point for its pass rate, cost and harness; click it for the model page. See how the frontier is computed →
Method
How it is computed
The number
cost per solved task = cost per attempt ÷ pass rate
What it costs to get a task finished, not what a token costs.
Measured
the benchmark observed the cost
No Solvency assumption is inside these figures. Tier and assumptions cannot move them.
Modelled
pass rate published, cost from verified prices
Solvency's loop model, labelled as an assumption and adjustable. Never ranked against measured rows.