Spokes.wiki Search About
Article source ↗ source url updated Tue Aug 04 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

LLM API Pricing Comparison (CloudZero, 2026)

A ranked cost comparison of the major LLM APIs — the core source for llm-api-pricing. Headline: a ~600× cost spread across the market, $/1M tokens (input/output).

Snapshot tiers (mid-2026; volatile)

  • Budget ($0.10–$0.75 in): Mistral Small 3.2 $0.10/$0.30 · GPT-4.1 Nano $0.10/$0.40 · deepseek V3.2 $0.27/$1.10.
  • Mid / production ($2–$5 in): GPT-5.4 $2.50/$15 · Claude Sonnet 4.6 $3/$15 · Gemini 3.1 Pro (≤200K) $2/$12.
  • Frontier / reasoning ($15–$30 in): GPT-5.4 Pro $30/$180 · o3 $15/$60 · Claude Opus 4.7 $5/$25.

Provider positioning

  • OpenAI — widest range (~150× spread), pick by workload.
  • Anthropic (anthropic) — consistency + aggressive prompt caching (90% off cached).
  • google — undercuts on sticker price (Gemini 2.5 Flash $0.15/$0.60 = “cheapest high-quality option from a major U.S. provider”).

Cost levers (stackable → ~25% of list)

Prompt caching 50–90% off repeated input · batch flat 50% off · model routing to budget models cuts 70–90% of qualifying calls. Structural fact: output tokens cost 2–6× input, so context optimization is the biggest lever.

llm-api-pricing · deepseek · anthropic · llm-provider

Re-verified 2026-08-04 — the prices are stale, and the corpus can prove it

Source re-read: published 2026-05-11, with no update since, listing GPT-5.4 at $2.50/$15, Claude Sonnet 4.6 at $3.00/$15, Claude Opus 4.7 at $5.00/$25, Gemini 2.5 Flash at $0.30/$2.50, and DeepSeek deepseek-chat at $0.27/$1.10 with deepseek-reasoner at $0.55/$2.19.

Two of those DeepSeek entries name models that no longer exist. deepseek-api-docs, re-read from the vendor’s own pricing page the same day, records deepseek-chat and deepseek-reasoner as retired on 2026-07-24, replaced by deepseek-v4-flash ($0.14/$0.28) and deepseek-v4-pro ($0.435/$0.87). The comparison is not merely old; it prices a discontinued product, and the replacement is half the input cost of what it lists.

The model lineup is dated on the other axis too: llm-leaderboard-stats, also re-read today, has Claude Opus 5 and GPT-5.6 at the top, neither of which this table contains.

Kept, not deleted, and the tier is unchanged: as a record of the May 2026 price landscape it is still usable, and the drift is now measured rather than suspected. Do not cite any number here as current. This is the clearest demonstration in the corpus of why the freshness gate exists — a third-party comparison went wrong within twelve weeks, and only a first-party check caught it.