Skip to content
SOLVENCY

Which model is cheapest for what you actually ship?

Cost per solved task — what it takes to get a task finished, not the token price. Every number carries a source and a verified date.

I build tasks,
about a month.

multi-file refactor, feature with tests · Tier: moderate — measured rows unchanged.

DeepSeek V4 Flash costs $0.12 per solved task against Claude Opus 5 at $12.01 — ▼ 100x cheaper for 18 fewer points of pass rate. Over 200 tasks that is $2.4k a month.

Skip the ranked rows

Measured · The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.

MEASURED: cost per solved task, ranked6 measured rows ranked cheapest first. 1 DeepSeek V4 Flash $0.12; 2 Gemini 3.7 Flash $2.12; 3 Grok 4.5 $3.81; 4 GPT-5.6 Sol $9.88; 5 Claude Opus 5 $12.01; 6 Claude Fable 5 $17.46. Monthly figures at 200 tasks.harness Codex · pass rate 50% · $0.06 per attemptDeepSeek V4 Flash$0.12vs ›$24/moharness Opencode · pass rate 60% · $1.27 per attemptGemini 3.7 Flash$2.12vs ›$423/moharness Grok Build · pass rate 64% · $2.44 per attemptGrok 4.5$3.81vs ›$763/moharness Codex · pass rate 65% · $6.42 per attemptGPT-5.6 Sol$9.88vs ›$2.0k/moharness Claude Code · pass rate 68% · $8.17 per attemptClaude Opus 5$12.01vs ›$2.4k/moharness Claude Code · pass rate 67% · $11.70 per attemptClaude Fable 5$17.46vs ›$3.5k/mo

Modelled · Pass rate published; cost is Solvency's loop model at verified prices.

MODELLED: cost per solved task, ranked4 modelled rows ranked cheapest first. 1 GPT-5.4 $1.86; 2 Gemini 3.1 Pro (preview) $1.91; 3 Claude Opus 4.6 $3.85; 4 Claude Opus 4.5 $4.36. Monthly figures at 200 tasks.pass rate 59% · $1.10 per attemptGPT-5.4$1.86vs ›$372/mopass rate 46% · $0.88 per attemptGemini 3.1 Pro (preview)$1.91vs ›$382/mopass rate 52% · $2.00 per attemptClaude Opus 4.6$3.85vs ›$771/mopass rate 46% · $2.00 per attemptClaude Opus 4.5$4.36vs ›$872/mo
Stale pass rates (7) · Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.
STALE: cost per solved task, ranked7 stale rows ranked cheapest first. 1 GPT-5 $0.74; 2 Gemini 2.5 Pro $0.78; 3 o3 $0.89; 4 GPT-4.1 $1.37; 5 Claude Sonnet 4 $1.96; 6 Claude Opus 4 $8.33; 7 o3-pro $8.48. Monthly figures at 200 tasks.pass rate 88% · $0.65 per attemptGPT-5$0.74vs ›$148/mopass rate 83% · $0.65 per attemptGemini 2.5 Pro$0.78vs ›$156/mopass rate 81% · $0.72 per attempto3$0.89vs ›$177/mopass rate 52% · $0.72 per attemptGPT-4.1$1.37vs ›$275/mopass rate 61% · $1.20 per attemptClaude Sonnet 4$1.96vs ›$392/mopass rate 72% · $6.00 per attemptClaude Opus 4$8.33vs ›$1.7k/mopass rate 85% · $7.20 per attempto3-pro$8.48vs ›$1.7k/mo
Not shown (8) · no published pass rate, reported as missing, never estimated

No published pass rate, so reported as missing rather than estimated: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.3 Codex, DeepSeek V4 Pro, Grok 4.6, Mistral Medium 3.5.

Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · Sources

Permalink with your name, re-priced when prices change.

Pro (soon): export CSV/JSON · price-change alerts · your own prices →

Assumptions (4) · move modelled rows only

These move modelled rows only. A model with no published cached-input price is computed uncached and says so in its row, rather than dropping out of its group. Measured rows carry a cost the benchmark observed and cannot be changed by any assumption. Frontier efficiency has no published source and is applied only to modelled rows — untick it to see unadjusted results. Methodology.

The frontier

Why not just pick the top score?

Lower and to the right is better. The line joins the measured models nobody beats on both axes.

Token price predicts almost nothing: DeepSeek V4 Flash is 19x cheaper per output token than Claude Opus 5 — and 100x cheaper per solved task.

Read the note →
Cost per solved task against pass rate, with the measured Pareto frontierScatter of 10 models: x is pass rate 40–80%, y is dollars per solved task on a log scale from $0.1 to $100. The frontier joins the measured models nobody beats on both axes: DeepSeek V4 Flash (50%, $0.12), Gemini 3.7 Flash (60%, $2.12), Grok 4.5 (64%, $3.81), GPT-5.6 Sol (65%, $9.88), Claude Opus 5 (68%, $12.01).$0.1$1$10$10040%50%60%70%80%PASS RATE →$ / SOLVED TASK (log) ↑DeepSeek V4 FlashGemini 3.7 FlashGrok 4.5GPT-5.6 SolClaude Opus 5Claude Fable 5

Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21

Modelled points move with your tier and assumptions; measured points never do. Hover or focus a point for its pass rate, cost and harness; click it for the model page. See how the frontier is computed →

Method

How it is computed

The number

cost per solved task = cost per attempt ÷ pass rate

What it costs to get a task finished, not what a token costs.

Measured

the benchmark observed the cost

No Solvency assumption is inside these figures. Tier and assumptions cannot move them.

Modelled

pass rate published, cost from verified prices

Solvency's loop model, labelled as an assumption and adjustable. Never ranked against measured rows.

Methodology · Sources · Research note 01