Spokes.wiki Search About
Article source ↗ source url updated Wed Jul 08 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Grok 4.5 — price vs. benchmark gaps (The Decoder)

The Decoder’s argument that xAI‘s Grok 4.5 is priced low enough that its weaker coding benchmarks may not matter for buyers who optimize cost-per-task rather than peak score.

Pricing (per 1M tokens, input / output — dated snapshot)

  • Grok 4.5 — $2 / $6
  • Opus 4.8 — $5 / $25
  • GPT-5.5/5.6 — $5 / $30
  • Fable 5 — $10 / $50

Grok 4.5 is the cheapest of the four on both input and output — roughly a third of Opus 4.8’s output rate and about an eighth of Fable 5’s.

Benchmarks (as reported)

ModelDeepSWE 1.1Terminal Bench 2.1SWE Bench Pro
Fable 570%84.3%80.4%
GPT 5.567%83.4%58.6%
Grok 4.553%83.3%64.7%
Opus 4.859%78.9%69.2%

Grok 4.5 trails clearly on DeepSWE (GitHub-issue resolution) but is competitive on Terminal Bench and mid-pack on SWE Bench Pro.

The argument

“Lower per-token pricing and fewer tokens per task” is the pitch. xAI claims Grok 4.5 uses 4.2× fewer tokens than Opus 4.8 on relevant tasks and runs at ~80 tokens/second, so real cost-per-output beats the sticker gap even where the benchmark trails. The Decoder frames this as the same play the Chinese labs run (deepseek, qwen, z-ai): be good enough, then win on price and token efficiency rather than topping the leaderboard — see llm-api-pricing and llm-benchmarks.

Tier

T3 — established AI-news outlet reporting specific prices and third-party benchmark figures, but not the primary rate cards; the token-efficiency numbers are xAI’s own claims. freshness: volatile — prices and model versions move fast; treat as a July-2026 snapshot.

xai-grok · claude-fable-5 · llm-api-pricing · llm-benchmarks · deepseek · llm-provider