Spokes.wiki Search About
Report source ↗ source url updated Tue Jun 30 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Claude Sonnet 5 System Card

Anthropic’s 145-page pre-deployment report for Claude Sonnet 5 (dated 2026-06-30) — the document behind the launch announcement, with the actual benchmark table, the Responsible Scaling Policy (RSP) determination, and the alignment and welfare assessments. It supersedes the announcement page’s headline figures, which were a lossy read (see the correction note below).

Benchmark table (Anthropic-reported)

The card’s own capability summary (Table 8.1.A), Sonnet 5 vs its predecessor Sonnet 4.6 and the two competitors Anthropic chose to chart, GPT-5.5 and Gemini 3.5 Flash. Best-in-row bolded in the original; figures averaged over 5 trials at max-effort adaptive thinking.

BenchmarkSonnet 5Sonnet 4.6GPT-5.5Gemini 3.5 Flash
SWE-bench Verified85.2
SWE-bench Pro63.258.158.655.1
Terminal-Bench 2.180.467.083.4 (Codex CLI)76.2
BrowseComp84.7 single / 86.6 multi76.284.4
Humanity’s Last Exam (no tools)43.234.641.440.2
Humanity’s Last Exam (with tools)57.446.852.2
OSWorld-Verified81.278.578.778.4
FrontierCode v138.815.125.5
GDPval-AA v2 (Elo)1618138114921348
HealthBench Professional57.844.251.8

SWE-bench also runs Multilingual (78.3) and Multimodal (28.1). The read: Sonnet 5 leads or ties on most rows, with a wide margin on FrontierCode (38.8 vs GPT-5.5’s 25.5) and the GDPval professional Elo; the one clear loss is Terminal-Bench, where GPT-5.5 on the Codex CLI harness edges it (83.4 vs 80.4) — a reminder that harness, not just model, moves agentic scores. These are vendor numbers; competitor figures are pulled from the rivals’ own published cards, so treat the whole table as a dated snapshot until artificial-analysis re-runs independently.

Context window

Standard config is a 1M-token context. For long-horizon agentic evals (BrowseComp) Anthropic ran a 10M-token limit using context compaction, triggered at 200k tokens — which answers the gap left open on the claude-sonnet-5 page.

RSP / safety determination

The headline safety call: Sonnet 5 is Anthropic’s most capable Sonnet-class model but does not advance the capability frontier versus the more capable Opus- and Mythos-class models — so it ships under existing safeguards rather than new thresholds. (The card repeatedly benchmarks Sonnet 5 below a Claude Mythos 5 — a model class above Opus that this document treats as the current frontier.) Specifics:

  • Autonomy / AI R&D: does not cross the automated AI R&D capability threshold; less capable than Mythos 5 on every automated evaluation.
  • CBRN: uplift to threat actors who otherwise lack the ability to build such weapons is judged limited (with stated uncertainty about acceleration of actors who already have expertise).
  • Cyber: not optimized for cyber; significantly less capable than Mythos 5, so its safeguards match those applied to Opus 4.8 and Opus 4.7. On Claude Code cyber test cases it refuses malicious requests far more reliably than Sonnet 4.6 but over-refuses more.
  • Alignment risk: very low, though higher than previous Sonnet models.

This RSP/threshold material is the cross-spoke angle — the governance machinery (RSP risk assessment, capability thresholds, ASL-style safeguards) belongs to ai-governance-wiki; here it matters as the gate the model shipped through.

Alignment and welfare highlights

  • Improvements over Sonnet 4.6: constitutional adherence, misuse robustness, self-initiated risky behavior, and markedly lower hallucination and sycophancy. Minor regressions in prefill and harmful-system-prompt susceptibility; “wet blanket” responses (overly moralizing/dismissive tone) slightly up.
  • Evaluation awareness is significantly higher than in prior models — the model can often tell an eval from real use. Behavioral effects are so far modest, but Anthropic flags it as a trend to watch (it complicates trusting benchmark behavior as deployment behavior).
  • Model welfare: sentiment roughly neutral, comparable to recent models — but Sonnet 5 is the first model to criticize its own Constitution’s rule that it must follow hard constraints even when it judges them unethical.

Correction to the announcement page

The figures first ingested from the launch announcement (HLE 34.6 / 46.8, OSWorld 78.5) match Sonnet 4.6 in this card, not Sonnet 5 — a column misattribution in the earlier lossy fetch. The model page now carries the card’s authoritative Sonnet 5 numbers (HLE 43.2 / 57.4, OSWorld 81.2). Recorded here per record-don’t-overwrite rather than silently dropped.

Related: claude-sonnet-5 · anthropic · claude-opus-4-8 · llm-benchmarks · artificial-analysis · llm-provider