Spokes.wiki Search About
Dataset source ↗ source url updated Tue Aug 04 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

LLM Stats Leaderboard

An independent leaderboard ranking 300+ models by a composite score blending verified benchmarks, live performance, and price — the core source for llm-benchmarks. Pricing refreshes hourly; live metrics update continuously.

What goes into the score

  • Benchmarks: GPQA Diamond, SWE-Bench Verified, MMLU, AIME, HumanEval, coding-arena.
  • Speed: output throughput (tokens/sec) + time-to-first-token, 7-day rolling.
  • Price: per-token input cost from official lists.
  • Context window and agentic/coding performance.

Top of the board — re-read 2026-08-04

#ModelScore
1Claude Opus 5 (anthropic)57.8
2GPT-5.6 Sol (OpenAI)57.7
3**[[claude-fable-5Claude Fable 5]]** (Anthropic)
4Claude Mythos Preview (Anthropic)56.1
5Kimi K3 (Moonshot AI)55.7
6GPT-5.6 Terra (OpenAI)53.1
7Qwen3.8 Max (Alibaba)52.6
8claude-opus-4-8 (Anthropic)52.5

What moved since the 2026-06-01 read, kept rather than overwritten because the churn is the point this page exists to make:

  • The old top two — Mythos Preview #1 and claude-opus-4-8 #2 — are now #4 and #8. Opus 4.8 fell six places in nine weeks without getting worse; the board moved around it.
  • GPT-5.5 was #3 and is off the top eight entirely, replaced by two GPT-5.6 variants.
  • Qwen3.7 MaxQwen3.8 Max, still the Alibaba entry in the same band.
  • The gap between #1 and #2 is 0.1 points (57.8 vs 57.7), and #1–#4 spans 1.7. Treat the ordering at the top as noise and the band as the signal.

The June read also recorded Grok 4 Fast on context (2.0M) and Mercury 2 on throughput (784 tok/s); those axes were not re-read this pass, so they stand as a June snapshot rather than a current one.

Why it matters here

It operationalizes “best model” as a multi-axis tradeoff (intelligence × speed × price × context), not a single number — the antidote to single-benchmark hype. Anthropic/OpenAI lead reasoning while Alibaba (Qwen) and xAI compete on cost/context. (Leaderboard = a dated snapshot; rankings churn weekly — and the 2026-06-01 → 2026-08-04 diff above is the evidence: the top two both fell, one of them by six places, with no change to the models themselves.)

llm-benchmarks · anthropic · claude-opus-4-8 · llm-provider · llm-api-pricing