Grok 4.5 — price vs. benchmark gaps (The Decoder)
The Decoder’s argument that xAI‘s Grok 4.5 is priced low enough that its weaker coding benchmarks may not matter for buyers who optimize cost-per-task rather than peak score.
Pricing (per 1M tokens, input / output — dated snapshot)
- Grok 4.5 — $2 / $6
- Opus 4.8 — $5 / $25
- GPT-5.5/5.6 — $5 / $30
- Fable 5 — $10 / $50
Grok 4.5 is the cheapest of the four on both input and output — roughly a third of Opus 4.8’s output rate and about an eighth of Fable 5’s.
Benchmarks (as reported)
| Model | DeepSWE 1.1 | Terminal Bench 2.1 | SWE Bench Pro |
|---|---|---|---|
| Fable 5 | 70% | 84.3% | 80.4% |
| GPT 5.5 | 67% | 83.4% | 58.6% |
| Grok 4.5 | 53% | 83.3% | 64.7% |
| Opus 4.8 | 59% | 78.9% | 69.2% |
Grok 4.5 trails clearly on DeepSWE (GitHub-issue resolution) but is competitive on Terminal Bench and mid-pack on SWE Bench Pro.
The argument
“Lower per-token pricing and fewer tokens per task” is the pitch. xAI claims Grok 4.5 uses 4.2× fewer tokens than Opus 4.8 on relevant tasks and runs at ~80 tokens/second, so real cost-per-output beats the sticker gap even where the benchmark trails. The Decoder frames this as the same play the Chinese labs run (deepseek, qwen, z-ai): be good enough, then win on price and token efficiency rather than topping the leaderboard — see llm-api-pricing and llm-benchmarks.
Tier
T3 — established AI-news outlet reporting specific prices and third-party benchmark figures, but not the
primary rate cards; the token-efficiency numbers are xAI’s own claims. freshness: volatile — prices and
model versions move fast; treat as a July-2026 snapshot.
Related
xai-grok · claude-fable-5 · llm-api-pricing · llm-benchmarks · deepseek · llm-provider