SOLVENCY BENCH · LONG-HORIZON · Solvency ran these itself: 30 hard agentic repo tasks, sandboxed tool loop, hidden tests (bench/ in the repo). Own population — the tier models are judged on.
SOLVENCY BENCH · LONG-HORIZON · Solvency ran these itself: 30 hard agentic repo tasks, sandboxed tool loop, hidden tests (bench/ in the repo). Own population — the tier models are judged on.
Token price predicts almost nothing: DeepSeek V4 Pro is 13x cheaper per output token than Claude Fable 5 — and 83x cheaper per solved task. Lower cost and higher pass rate is better; the purple line joins the measured models nobody beats on both.
Create a free account to keep assumptions and save this scenario.
GitHub, Google or email. Assumptions move modeled rows only; measured rows never change.
Priced, awaiting measurement (292) · verified prices for every listed model; cost per solved task stays missing until a benchmark measures it — shown in full, never estimated.
These move modeled rows only. A model with no published cached-input price is computed uncached and says so in its row, rather than dropping out of its group. Measured rows carry a cost the benchmark observed and cannot be changed by any assumption. Frontier efficiency has no published source and is applied only to modeled rows — untick it to see unadjusted results. Methodology.