Independent cost data · 6 measured models · verified 2026-08-21
Token price is not task cost.
Pricing pages compare dollars per million tokens. Solvency ranks AI coding models by cost per solved task — what it costs to get a task finished — with a source and a verification date on every number.
Per input token
11x
DeepSeek V4 Flash vs Claude Opus 5, the top-scoring model
Per output token
19x
cheaper
Per solved task
100x
$0.12 vs $12.01
Spread, cheapest to dearest
38x → 146x
token price → cost per solved task
Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · measured on agentic coding tasks; prices from provider pricing pages
Cost per solved task
measuredUSD per task that passes · lower is better
Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21
Full leaderboard ↓Leaderboard
Cost per solved task, ranked
Cost per attempt divided by pass rate. Lower is better. Measured and modelled rows are ranked separately and never averaged together.
Measured · cost per solved task
The benchmark ran the model and observed the cost. No Solvency assumption is inside these figures.
| # | Model | Harness | Index | $ / task | $ / solved task |
|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash | Codex | 50 | $0.06 | $0.12 |
| 2 | Gemini 3.7 Flash | Opencode | 60 | $1.27 | $2.12 |
| 3 | Grok 4.5 | Grok Build | 64 | $2.44 | $3.81 |
| 4 | GPT-5.6 Sol | Codex | 65 | $6.42 | $9.88 |
| 5 | Claude Opus 5 | Claude Code | 68 | $8.17 | $12.01 |
| 6 | Claude Fable 5 | Claude Code | 67 | $11.70 | $17.46 |
Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · Coding Agent Index v1.4, 326 tasks, 3 attempts each. Rows are harness + model pairs, not bare models. $ / solved task = $ / task ÷ (index ÷ 100).
Modelled · pass rate published, cost estimated
These sources publish no cost. Cost is Solvency's loop model at today's verified prices — an assumption, labelled as one. Not comparable with the measured table.
| # | Model | Source | Pass rate from | Pass | $ / solved task |
|---|---|---|---|---|---|
| 1 | GPT-5.4 | SEAL | 2026-08-21 | 59% | $4.19 |
| 2 | Gemini 3.1 Pro (preview) | SEAL | 2026-08-21 | 46% | $4.30 |
| 3 | Claude Opus 4.6 | SEAL | 2026-08-21 | 52% | $8.67 |
| 4 | Claude Opus 4.5 | SEAL | 2026-08-21 | 46% | $9.81 |
| 5 | GPT-5 | Aider | 2025-08-23 · stale | 88% | $1.66 |
| 6 | Gemini 2.5 Pro | Aider | 2025-06-06 · stale | 83% | $1.76 |
| 7 | o3 | Aider | 2025-06-25 · stale | 81% | $1.99 |
| 8 | Claude Sonnet 4 | Aider | 2025-05-24 · stale | 61% | $4.40 |
| 9 | GPT-4.1 | Aider | 2025-04-14 · stale | 52% | $5.15 |
| 10 | Claude Opus 4 | Aider | 2025-05-25 · stale | 72% | $18.75 |
| 11 | o3-pro | Aider | 2025-06-28 · stale | 85% | $19.08 |
Sources: Scale SEAL (SWE-bench Pro) and Aider polyglot · verified 2026-08-21 · Aider pass rates predate 2026 and are marked stale. Loop model and frontier-efficiency assumptions are listed in the methodology.
Not ranked — no published pass rate, reported as missing rather than estimated: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.3 Codex, DeepSeek V4 Pro, Grok 4.6, Mistral Medium 3.5.
Research note 01 · 2026-08-21
Cost Per Solved Task
Per-token pricing does not predict what a coding task actually costs. Measured across six current models, the gap widens from 38x to 146x.
Read the note →
Changelog
2026-08-21 · initial
Initial dataset published: 25 models priced, each verified against its provider's own pricing page.
Highlights
Three ways to rank the same six models
Same measured rows, three axes. The order on the left is the one that predicts your bill; the order on the right is the one pricing pages publish.
Cost per solved task
USD per task that passes · lower is better
Coding Agent Index
Composite score, 0–100 · higher is better
Output price
USD per million output tokens · what pricing pages show
Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · index and per-task cost; output prices from each provider's pricing page
Calculator
Your task mix, your volume
Pick a task tier and a monthly volume. Measured rows cannot be moved by the assumptions; modelled rows can, and say so.
Sign in to run your own scenario
Free account. Set your task tier, volume, cache rate and takeover cost; share the result by link.
Measured
The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.
| # | Model | Pass | $ / solved task | $ / month |
|---|---|---|---|---|
| 1 | DeepSeek V4 Flash measured | 50% | $0.12 | $24.00 |
| 2 | Gemini 3.7 Flash measured | 60% | $2.12 | $423 |
| 3 | Grok 4.5 measured | 64% | $3.81 | $763 |
| 4 | GPT-5.6 Sol measured | 65% | $9.88 | $2.0k |
| 5 | Claude Opus 5 measured | 68% | $12.01 | $2.4k |
| 6 | Claude Fable 5 measured | 67% | $17.46 | $3.5k |
Modelled
Pass rate published, cost estimated by Solvency's loop model — an assumption.
| # | Model | Pass | $ / solved task | $ / month |
|---|---|---|---|---|
| 1 | GPT-5.4 modelled | 59% | $1.86 | $372 |
| 2 | Gemini 3.1 Pro (preview) modelled | 46% | $1.91 | $382 |
| 3 | Claude Opus 4.6 modelled | 52% | $3.85 | $771 |
| 4 | Claude Opus 4.5 modelled | 46% | $4.36 | $872 |
Modelled from stale pass rates
Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.
| # | Model | Pass | $ / solved task | $ / month |
|---|---|---|---|---|
| 1 | GPT-5 stale | 88% | $0.74 | $148 |
| 2 | Gemini 2.5 Pro stale | 83% | $0.78 | $156 |
| 3 | o3 stale | 81% | $0.89 | $177 |
| 4 | GPT-4.1 stale | 52% | $1.37 | $275 |
| 5 | Claude Sonnet 4 stale | 61% | $1.96 | $392 |
| 6 | Claude Opus 4 stale | 72% | $8.33 | $1.7k |
| 7 | o3-pro stale | 85% | $8.48 | $1.7k |
Measured rows carry a cost the benchmark observed, so the tier, cache and efficiency controls cannot move them. Modelled rows are priced by an assumed loop model.
Finding 1
Token price is not task cost
Log scale, normalised to the cheapest model on each axis. Measured rows only, so no Solvency assumption is inside these numbers.
Cost per solved task = measured cost per task / pass rate. Source: Artificial Analysis (artificialanalysis.ai) — Coding Agent Index v1.4. Prices verified 2026-08-21.
Finding 2
What one index point costs
Coding Agent Index vs cost per solved task. Dashed line: the frontier — nothing cheaper scores higher.
Rows are harness+model pairs, not bare models. Source: Artificial Analysis (artificialanalysis.ai) — Coding Agent Index v1.4. Verified 2026-08-21.
Finding 3
Newest entry, by source
Bar runs from each source’s newest published entry to today. Longer is worse.
Aider verified against its raw leaderboard YAML. SEAL publishes no update date, so its staleness is unknown, not zero.
Sources
Where every number comes from
In preference order: fewest Solvency assumptions first, then freshness. Benchmark data is cited and linked, never redistributed.
| Source | Tasks | Covers 2026 models | Publishes cost | Basis | Newest entry | Verified |
|---|---|---|---|---|---|---|
| Artificial Analysis Coding Agent Index v1.4 | 326 | yes | yes, measured | measured by source | 2026-08-21 | 2026-08-21 |
| Scale SEAL leaderboard - SWE-bench Pro (public) | 1,865 | yes | not captured | modelled by Solvency | unknown | 2026-08-21 |
| Aider polyglot benchmark | 225 | no | historical only | historical at run date | 2025-10-03 | 2026-08-21 |
· Prices are verified against each provider's own pricing page; recalled prices are never used. Missing is printed as missing.
Share the finding
The tweet text is generated from the data, so it matches the tables above.