Introducing Claude Opus 5 (Anthropic)
Anthropic’s launch post for Claude Opus 5, published 2026-07-24. First-party vendor announcement, so T1 for provenance and T3-ish for the benchmark claims — every number below is self-reported by the lab selling the model, and the comparison set is chosen by Anthropic.
What it says
The pitch is price/performance, not a new ceiling. Opus 5 “comes close to the frontier intelligence of Claude Fable 5 at half the price.” That framing matters for this spoke: Anthropic is positioning its flagship below Fable 5 on capability and explicitly selling the gap as a discount, which is the first T1 confirmation of the ordering three third-party sources had already guessed at (see claude-fable-5).
Price: $5 / $25 per million input/output tokens — identical to Opus 4.8. A generation of capability at a flat sticker. There is also a fast mode at 2× the base price running “around 2.5 times the default speed,” which is a new axis for this spoke: speed sold as a paid tier on the same model rather than as a separate cheaper model. (Compare cerebras-pricing, where speed is the whole product.)
Effort still gates cost. The post repeats the Sonnet-5-era effort lever with the level set
written out as low, high, xhigh, and max — customers “optimize for intelligence or conserve
tokens.” Rate alone still doesn’t set the bill (see llm-api-pricing).
That list is incomplete, and the omission matters (checked 2026-08-04 against Anthropic’s own docs, which say verbatim: “Claude Opus 5 supports all five effort levels.”). The missing level is
medium— the one the same documentation tells developers to use “liberally” alongsidelowas the primary control for token cost and response time. Quoted here as the post has it, per record-don’t-overwrite; the corrected ladder is on claude-opus-5. Worth noting which direction the loss runs: the marketing copy dropped a cost-saving option, not a capability claim.
Benchmarks, all vendor-reported and mostly stated as ratios rather than absolute scores:
| Benchmark | Claim |
|---|---|
| Frontier-Bench v0.1 | ”surpasses all other models,” and more than doubles Opus 4.8 at a lower cost per task |
| CursorBench 3.2 (max effort) | within 0.5% of Fable 5’s peak, at half the cost per task |
| ARC-AGI 3 | 3× the next-best model |
| Zapier AutomationBench | pass rate ~1.5× the next-best model |
| OSWorld 2.0 | beats Fable 5’s best result at just over a third of the cost |
| Organic chemistry | +10.2 points over Opus 4.8 |
With no absolute scores and no competitor named per row, “next-best model” can’t be checked against anything. Five of the six claims are framed as cost per task rather than as a score.
Availability: Claude API, Claude.ai, Claude Code, and Claude Cowork, plus the reseller platforms (Claude on AWS, Google Cloud Vertex AI, Microsoft Foundry) — the post names the partners but not the per-model rollout. Opus 4.8 is not deprecated; it stays as the fallback tier.
Safety. Anthropic calls it its “most aligned model to date” on an automated behavioral audit, with Opus-4.8-like safeguards plus “stronger guardrails on a narrow range of cyber tasks.” Flagged cyber requests in Claude Cowork fall back to Opus 4.8 by default — a routing behaviour, not a refusal (cf. claude-refusals-and-fallback).
Mythos 5 named at last. The Sonnet 5 system card revealed a class above Opus called Mythos; here Claude Mythos 5 appears as a shipping comparison point. It is “substantially behind” Opus 5 on most benchmarks but stays ahead on cybersecurity tasks and biology research — so Mythos reads as a narrow, domain-strong line, not the general ceiling the earlier system-card mention implied.
What’s missing
No context window is stated anywhere in the post — unusual for a flagship launch, and a gap against Sonnet 5’s advertised 1M. No absolute benchmark scores, no per-row competitor names, and no system card linked yet. Treat the whole table as a dated vendor snapshot.
Related
claude-opus-5 · claude-fable-5 · claude-sonnet-5 · llm-api-pricing · llm-benchmarks · anthropic · claude-opus-4-8