Data snapshot · 15 Aug 2026
Fifteen models. Five charts. Benchmarks, pricing, context and speed — pulled from BenchLM, Artificial Analysis, Vals AI and vendor announcements. Everything you need to pick a model, argue about models, or build with the right one.
01 — Raw capability
BenchLM Capability index, 0–100, 218 ranked models, field median 58.2. The top four are separated by 1.2 points. The open-weight champ is Qwen3.8 Max at #6.
higher = more capable · median 58.2
02 — Coding
Real GitHub issues, real fixes. DeepSeek V4 Pro 0813 sits 0.6 points behind Claude Opus 5 while costing roughly 1/30th as much.
Vals AI leaderboard · captured scores only
agentic terminal work
| Model | Score |
|---|---|
| GPT-5.6 Sol | 91.9% |
| Claude Mythos 5 | 88.0% |
| GPT-5.6 Terra | 87.4% |
MMLU-Pro · LiveCodeBench · ARC-AGI-2
| Benchmark | Leader |
|---|---|
| MMLU-Pro | Qwen3.7 Max 89.6% |
| LiveCodeBench | Gemini 3 Pro 91.7% |
| ARC-AGI-2 | Gemini 3.1 Pro 77.1% |
03 — Money
Output price, USD per 1M tokens. Frontier prices are 88% below March 2023 levels — but the gap between $0.79 and $50 is the whole game.
OpenRouter / vendor list prices · base tiers
04 — Value
Capability vs output price, log scale. Bottom-right quadrant = frontier capability at commodity prices. That is DeepSeek's whole argument.
BenchLM capability · $ per 1M output · log x-axis
#1 AA Intelligence Index (63.0), #1 SWE-bench (97.0%), #1 agentic. 1M context. $5/$25. The current champion, full stop.
Opus 5 owns SWE-bench. Sol leads Terminal-Bench 2.0 (91.9%) and the AA Coding Agent Index (80) — near-Fable-5 intelligence at ~⅓ the cost.
96.4% SWE-bench at $0.40/$0.79. MIT open weights, 1M context, 90 tok/s. 0.6 points off #1 at ~1/30th the price.
#6 overall on BenchLM (79.9), #1 reasoning category. Highest-ranked open-weight model. Self-host economics.
27B dense, 262K context, multimodal. "Punches above its weight" — but needs a 24GB GPU at Q4. 16GB machines should stay at 14B-class models.
1M context, 113 tok/s, $2/$12. ARC-AGI-2 77.1% — screenshots, documents, charts, grounded reasoning. Trails the frontier on pure coding.
05 — Speed
measured values · BenchLM / Artificial Analysis
06 — Full table
| Model | Maker | Released | Open | Context | In $/M | Out $/M | Speed | SWE-bench V | BenchLM | AA Index |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 24 Jul 26 | ✗ | 1.0M | 5.00 | 25.00 | 53 | 97.0% | 83.1 #2 | 63.0 #1 |
| Claude Mythos 5 | Anthropic | 9 Jun 26 | ✗ restricted | 1.0M | 10.00 | 50.00 | n/m | — | 83.2 #1 | — |
| Claude Fable 5 | Anthropic | 9 Jun 26 | ✗ | 1.0M | 10.00 | 50.00 | 63 | — | 83.0 #3 | 62.1 #2 |
| Claude Opus 4.6 | Anthropic | 5 Feb 26 | ✗ | 1.0M | 5.00 | 25.00 | 39 | — | 68.0 #21 | — |
| GPT-5.6 Sol | OpenAI | 9 Jul 26 | ✗ | 1.05M | 5.00 | 30.00 | 62 | — | 82.0 #4 | 59 |
| GPT-5.3 Codex | OpenAI | 5 Feb 26 | ✗ | 400K | 1.75 | 14.00 | 111 | — | 65.9 #31 | — |
| Gemini 3.1 Pro | 19 Feb 26 | ✗ | 1.0M | 2.00 | 12.00 | 113 | — | 56.0 #92 | — | |
| Grok 4.6 | xAI | 12 Aug 26 | ✗ | 500K | 2.00 | 6.00 | 66 | — | 63.4 #43 | 60.9 #3 |
| DeepSeek V4 Pro 0813 | DeepSeek | 13 Aug 26 | ✓ MIT | 1.0M | 0.40 | 0.79 | 90 | 96.4% | 61.2 #52 | 53 |
| Kimi K3 | Moonshot | 16 Jul 26 | ✓ | 1.0M | 2.80 | 14.00 | — | 93.4% | — | — |
| Qwen3.8 Max | Alibaba | 3 Aug 26 | ✓ | 1.0M | n/a | n/a | 47 | — | 79.9 #6 | — |
| Qwen3.8-27B | Alibaba | Aug 26 | ✓ | 262K | n/a | n/a | n/m | — | unranked | — |
| GLM-5.3 | Zhipu Z.ai | 14 Aug 26 | ✓ MIT* | 1.0M | n/a* | n/a* | n/m | — | unranked | — |
| MiniMax M3 | MiniMax | 31 May 26 | ✓ | 1.0M | 0.23 | 0.96 | 100 | — | — | 45 |
| Llama 4 Maverick | Meta | 28 Feb 26 | ✓ | 1.0M | n/a | n/a | 119 | — | 22.8 #208 | — |
*GLM-5.3 pricing pending, MIT weights ~28 Aug. n/m not measured. n/a not published (self-host). Opus 4.8 88.6% and Grok 4.5 86.6% SWE-bench appear in the chart only.