LLM Stats Leaderboard
An independent leaderboard ranking 300+ models by a composite score blending verified benchmarks, live performance, and price — the core source for llm-benchmarks. Pricing refreshes hourly; live metrics update continuously.
What goes into the score
- Benchmarks: GPQA Diamond, SWE-Bench Verified, MMLU, AIME, HumanEval, coding-arena.
- Speed: output throughput (tokens/sec) + time-to-first-token, 7-day rolling.
- Price: per-token input cost from official lists.
- Context window and agentic/coding performance.
Top of the board — re-read 2026-08-04
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (anthropic) | 57.8 |
| 2 | GPT-5.6 Sol (OpenAI) | 57.7 |
| 3 | **[[claude-fable-5 | Claude Fable 5]]** (Anthropic) |
| 4 | Claude Mythos Preview (Anthropic) | 56.1 |
| 5 | Kimi K3 (Moonshot AI) | 55.7 |
| 6 | GPT-5.6 Terra (OpenAI) | 53.1 |
| 7 | Qwen3.8 Max (Alibaba) | 52.6 |
| 8 | claude-opus-4-8 (Anthropic) | 52.5 |
What moved since the 2026-06-01 read, kept rather than overwritten because the churn is the point this page exists to make:
- The old top two — Mythos Preview #1 and claude-opus-4-8 #2 — are now #4 and #8. Opus 4.8 fell six places in nine weeks without getting worse; the board moved around it.
- GPT-5.5 was #3 and is off the top eight entirely, replaced by two GPT-5.6 variants.
- Qwen3.7 Max → Qwen3.8 Max, still the Alibaba entry in the same band.
- The gap between #1 and #2 is 0.1 points (57.8 vs 57.7), and #1–#4 spans 1.7. Treat the ordering at the top as noise and the band as the signal.
The June read also recorded Grok 4 Fast on context (2.0M) and Mercury 2 on throughput (784 tok/s); those axes were not re-read this pass, so they stand as a June snapshot rather than a current one.
Why it matters here
It operationalizes “best model” as a multi-axis tradeoff (intelligence × speed × price × context), not a single number — the antidote to single-benchmark hype. Anthropic/OpenAI lead reasoning while Alibaba (Qwen) and xAI compete on cost/context. (Leaderboard = a dated snapshot; rankings churn weekly — and the 2026-06-01 → 2026-08-04 diff above is the evidence: the top two both fell, one of them by six places, with no change to the models themselves.)
Related
llm-benchmarks · anthropic · claude-opus-4-8 · llm-provider · llm-api-pricing