Data snapshot · 15 Aug 2026

The frontier, charted.

Fifteen models. Five charts. Benchmarks, pricing, context and speed — pulled from BenchLM, Artificial Analysis, Vals AI and vendor announcements. Everything you need to pick a model, argue about models, or build with the right one.

Sources: BenchLM · Artificial Analysis · Vals AI · OpenRouter · vendor blogs  |  compiled 16 Aug 2026

01 — Raw capability

Who is actually the smartest?

BenchLM Capability index, 0–100, 218 ranked models, field median 58.2. The top four are separated by 1.2 points. The open-weight champ is Qwen3.8 Max at #6.

BenchLM Capability score

higher = more capable · median 58.2

02 — Coding

SWE-bench Verified. The score everyone argues about.

Real GitHub issues, real fixes. DeepSeek V4 Pro 0813 sits 0.6 points behind Claude Opus 5 while costing roughly 1/30th as much.

SWE-bench Verified (%)

Vals AI leaderboard · captured scores only

Terminal-Bench 2.0 leaderboard

agentic terminal work

ModelScore
GPT-5.6 Sol91.9%
Claude Mythos 588.0%
GPT-5.6 Terra87.4%

Other leaderboard tops

MMLU-Pro · LiveCodeBench · ARC-AGI-2

BenchmarkLeader
MMLU-ProQwen3.7 Max 89.6%
LiveCodeBenchGemini 3 Pro 91.7%
ARC-AGI-2Gemini 3.1 Pro 77.1%

03 — Money

What a million tokens costs.

Output price, USD per 1M tokens. Frontier prices are 88% below March 2023 levels — but the gap between $0.79 and $50 is the whole game.

Output price per 1M tokens ($, log scale)

OpenRouter / vendor list prices · base tiers

−88%frontier token prices vs Mar 2023
$1.00 / $3.75median API price, 145 models
−78%open-weight median vs proprietary
4773×most expensive vs cheapest

04 — Value

The value map. Where you'd be stupid not to.

Capability vs output price, log scale. Bottom-right quadrant = frontier capability at commodity prices. That is DeepSeek's whole argument.

Capability vs price

BenchLM capability · $ per 1M output · log x-axis

Best overall

Claude Opus 5

#1 AA Intelligence Index (63.0), #1 SWE-bench (97.0%), #1 agentic. 1M context. $5/$25. The current champion, full stop.

Best for coding

Claude Opus 5 / GPT-5.6 Sol

Opus 5 owns SWE-bench. Sol leads Terminal-Bench 2.0 (91.9%) and the AA Coding Agent Index (80) — near-Fable-5 intelligence at ~⅓ the cost.

Best value

DeepSeek V4 Pro 0813

96.4% SWE-bench at $0.40/$0.79. MIT open weights, 1M context, 90 tok/s. 0.6 points off #1 at ~1/30th the price.

Best open model

Qwen3.8 Max

#6 overall on BenchLM (79.9), #1 reasoning category. Highest-ranked open-weight model. Self-host economics.

Best local pick

Qwen3.8-27B

27B dense, 262K context, multimodal. "Punches above its weight" — but needs a 24GB GPU at Q4. 16GB machines should stay at 14B-class models.

Best multimodal

Gemini 3.1 Pro

1M context, 113 tok/s, $2/$12. ARC-AGI-2 77.1% — screenshots, documents, charts, grounded reasoning. Trails the frontier on pure coding.

05 — Speed

Tokens per second, because nobody likes waiting.

Output speed (tok/s)

measured values · BenchLM / Artificial Analysis

06 — Full table

Every number, in one place.

ModelMakerReleasedOpenContextIn $/MOut $/MSpeedSWE-bench VBenchLMAA Index
Claude Opus 5Anthropic24 Jul 261.0M5.0025.005397.0%83.1 #263.0 #1
Claude Mythos 5Anthropic9 Jun 26✗ restricted1.0M10.0050.00n/m83.2 #1
Claude Fable 5Anthropic9 Jun 261.0M10.0050.006383.0 #362.1 #2
Claude Opus 4.6Anthropic5 Feb 261.0M5.0025.003968.0 #21
GPT-5.6 SolOpenAI9 Jul 261.05M5.0030.006282.0 #459
GPT-5.3 CodexOpenAI5 Feb 26400K1.7514.0011165.9 #31
Gemini 3.1 ProGoogle19 Feb 261.0M2.0012.0011356.0 #92
Grok 4.6xAI12 Aug 26500K2.006.006663.4 #4360.9 #3
DeepSeek V4 Pro 0813DeepSeek13 Aug 26✓ MIT1.0M0.400.799096.4%61.2 #5253
Kimi K3Moonshot16 Jul 261.0M2.8014.0093.4%
Qwen3.8 MaxAlibaba3 Aug 261.0Mn/an/a4779.9 #6
Qwen3.8-27BAlibabaAug 26262Kn/an/an/munranked
GLM-5.3Zhipu Z.ai14 Aug 26✓ MIT*1.0Mn/a*n/a*n/munranked
MiniMax M3MiniMax31 May 261.0M0.230.9610045
Llama 4 MaverickMeta28 Feb 261.0Mn/an/a11922.8 #208

*GLM-5.3 pricing pending, MIT weights ~28 Aug. n/m not measured. n/a not published (self-host). Opus 4.8 88.6% and Grok 4.5 86.6% SWE-bench appear in the chart only.