Spokes.wiki Search About
Web Page source ↗ source url updated Sun Jul 19 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Cerebras pricing

Cerebras’ own pricing page for its inference API and coding subscriptions. Cerebras is not a model maker — it is a speed-differentiated inference provider that serves open models (Llama, Qwen and the like) on wafer-scale hardware cerebras-inference, so its pricing sells access + throughput, not a proprietary model. The page leads on the speed pitch (“20× faster than OpenAI and Anthropic”) rather than on a per-token rate card — and notably does not publish per-million-token prices for individual models, which is where most of the per-token market competes.

API tiers (mid-2026 snapshot — volatile)

  • Free trial — $5 in credits on signup; all models; community (Discord) support.
  • Developer — self-serve, from $10; “10× higher rate limits than free”; priority processing.
  • Enterprise — contact sales; highest rate limits, dedicated queue priority, custom model weights, fine-tuning/training, dedicated support.

Cerebras Code subscriptions (coding-agent product)

Flat monthly plans metered by a daily token bucket, not per-token billing:

  • Pro — $50/mo — up to 24M tokens/day (the page frames this as ~$48/day of value); aimed at indie/weekend developers. Listed sold out.
  • Max — $200/mo — up to 120M tokens/day (~$240/day of value); aimed at full-time and multi-agent use. Listed sold out.

Both coding tiers being sold out is itself a signal: wafer-scale capacity is scarce, so Cerebras rations its cheapest fast-inference access — consistent with speed, not price, being the thing it actually sells.

Access routes

Also reachable via AWS, OpenRouter, Hugging Face, and Vercel (aggregators/resellers), so Cerebras throughput shows up inside other platforms’ catalogs too.

Provenance, tier & cross-spoke

First-party vendor pricing page — T2, and a dated snapshot: tiers, dollar figures, and “sold out” status all drift; the pricing model (subscription token-buckets + per-use) is the durable part. Cross-spoke: the Cerebras Code subscriptions are a coding-agent product (an agentic-tooling-wiki concern) — noted here as context, not routed away; the wafer-scale mechanism and the company sit in llm-inference-wiki (cerebras-inference, cerebras-systems). See llm-api-pricing for how this speed-priced, no-rate-card model contrasts with the per-token market.