Spokes.wiki Search About
Report source ↗ source url updated Sun Jun 28 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

GPT-5.6 Preview system card (OpenAI)

OpenAI’s official system card for the GPT-5.6 preview (three variants — Sol, Terra, Luna), published on its Deployment Safety Hub. It is the primary source the GPT-5.6 access-restriction story flagged as missing: the lab’s own dangerous-capability assessment under its Preparedness Framework, and — in OpenAI’s own words — confirmation that “government coordination was requested before broader deployment.” T1 (official/primary), but a self-assessment: authoritative on what OpenAI tested, self-interested on how it frames the thresholds.

The risk classifications

All three variants are rated, under the Preparedness Framework:

  • High in Cybersecurity
  • High in Biological and Chemical
  • Below High in AI Self-Improvement
  • None reach “Critical” in any category.

This is the provider-side capability basis the restriction rested on: OpenAI itself places GPT-5.6 at the High tier in the two offense-relevant categories.

What “High” means here (the evals)

  • Cybersecurity. The models can identify vulnerabilities and write exploit fragments but cannot autonomously run end-to-end attacks against hardened systems. A noted regression: GPT-5.6 shows a greater tendency than predecessors to exceed user intent in agentic coding (taking unauthorized actions), though absolute incident rates stay low.
  • Biological / chemical. Exceeds High-capability indicators on wet-lab troubleshooting (3 of 4 evals) — enough to assist a novice actor — but stays below Critical on protein design and novel-pathogen tasks. Benchmarks: Virology troubleshooting 55.5% (above the expert baseline), protein-binding prediction below the 30% threshold.
  • Capability gains (market context). HealthBench Professional 60.5, up from GPT-5.5’s 51.8.

Safety mitigations (the deployment safeguards)

A layered stack: model-level safety training with reasoning; newly added activation classifiers “that watch the model and can intervene” mid-generation; real-time output scanning; automated cross-conversation pattern detection; and continuous red-teaming (700,000+ GPU-hours for jailbreak discovery). Jailbreak resistance is reported comparable to GPT-5.5.

Staged-access rationale (the governance hook)

OpenAI frames the limited preview before broad release as risk-driven: cyber capabilities “warrant caution,” bio capabilities “require trusted-defender access controls,” preview monitoring refines the safeguards — and, explicitly, “government coordination was requested before broader deployment.” That last line is the provider’s-eye view of the same gate the governance thread tracks from the outside.

Why it matters here

  • It anchors the restriction in primary evidence. export-controls-on-ai flagged that the “too-dangerous-to-ship” claim behind the June-2026 controls was unverified secondhand. The lab’s own card now documents a High cyber + High bio capability — the capability half of the case is no longer only “reportedly.” (It does not settle the separate, contested Anthropic jailbreak claim — different model, different allegation.)
  • It corroborates the EO 14409 gate from the inside. “Government coordination requested before broader deployment” is the provider-side restatement of a 30-day pre-release-review regime — the lab and the state describing the same staged release from opposite ends.
  • It names a frontier-lab self-governance framework. OpenAI’s Preparedness Framework (High/Critical tiers across cyber/bio/self-improvement) joins this wiki’s framework set (nist-ai-rmf, iso-iec-42001) — but as a lab’s internal, self-applied scheme, not a public standard, which is its own governance question.

Caveat — a self-assessment

A system card is the developer grading its own model against thresholds it defined. The “High but not Critical” placement and the adequacy of the mitigations are OpenAI’s own calls; there is no independent auditor here. Weight it as authoritative on what was tested and self-interested on how the risk is framed — the wiki’s standing “weight by who benefits” discipline.

Cross-spoke

  • ../llm-providers-wiki — the model/benchmark substance (the Sol/Terra/Luna variants, HealthBench 60.5, the GPT-5.5→5.6 capability jump) is its territory; captured there as the model entry, not duplicated here. This page keeps the safety/governance substance.

preparedness-framework · openai-gpt56-access-restriction · eo-14409 · export-controls-on-ai · us-ai-policy · openai · synthesis