MODEL

GPT-5.6 Sol

modeltopic-noteopenai

Overview

GPT-5.6 Sol is OpenAI‘s June 26, 2026 frontier model launching under US-government-approved access — the second wave of frontier releases gated under Trump’s June 2 frontier-AI EO and the subsequent Commerce Department directive (the first being Anthropic‘s Mythos). Sol benchmarks at 88.8% on Terminal-Bench 2.1, within-error parity with Mythos 5 at 88.0%, and prices at $5/$30 per million input/output tokens for the base tier. Sol ships alongside the broader GPT-5.6 family (Simon Willison‘s pricing read): Terra at $2.50/$15 (half the GPT-5.5 price) and Luna at $1/$6 as a new cheap tier.

Timeline

  • 2026-06-27-AI-DigestOpenAI releases GPT-5.6 Sol under the same US-government-approved access regime that already gated Anthropic‘s Mythos and Fable — Trump’s June 2 frontier-AI EO and the Commerce Department directive are the framing layer, and Sol’s launch is the second wave under that regime, not the start of a new one. Benchmark: 88.8% on Terminal-Bench 2.1 vs Mythos 5’s 88.0% (within-error tie). Pricing $5/$30 per M tokens base tier; Simon Willison surfaces the rest of the GPT-5.6 family alongside Sol (Terra at $2.50/$15, Luna at $1/$6). The Decoder reports OpenAI explicitly told government interlocutors the model is “not a preferred long-term model” for licensing of this kind (Decoder’s phrasing, not a direct Altman quote). Two reads: a competitive flagship lands at parity with Mythos at one-third the rumored deployment cost; the policy stack from MIT TR’s Anthropic-vs-government piece now reaches its second-lab consequence.
  • 2026-06-28-AI-DigestA clarifying detail surfaces on Sol’s gating: the “approving access customer by customer during this preview period” language is from a Sam Altman internal memo dated June 25, not Bloomberg or TechCrunch paraphrase, with the requesting bodies named as the Office of National Cyber Director plus OSTP. The framing the corpus now carries: OpenAI explicitly told government interlocutors “we don’t believe this kind of government access process should become the long-term default” — a publicly-recorded resistance to the very pattern Sol is being released under. Pair with today’s Mythos trusted-partner restoration: same Commerce-Department mechanism binds both labs, but Anthropic accommodates the pattern as a path back to deployment while OpenAI accommodates it under public objection. The 60-day test is whether OpenAI’s objection survives the next negotiated re-licensing or gets quietly absorbed into the regime.
  • 2026-07-03-AI-DigestOpenAI opens a limited preview of the full GPT-5.6 family to roughly 20 partner organisations (US government included) split across three tiers: Sol as flagship at $5/$30, Terra at $2.50/$15 (roughly 2× cheaper than GPT-5.5), Luna at $1/$6 as the low-cost tier — standing rates, not intro promos, with GA guided “in the coming weeks.” The practitioner-relevant lever change: new prompt-cache breakpoints with 30-minute minimum cache life, 1.25× cache-write premium, and 90% cache-read discount — long-lived agent scaffolds that stage large system prompts once get materially cheaper per additional turn than any prior OpenAI SKU. The three-tier shape mirrors Anthropic‘s Opus/Sonnet/Haiku split; cache mechanics target the same fat-system-prompt agent scaffold workload. Structural read the digest carries: labs now compete on standing base rates + cache economics rather than headline per-token cuts — the effective-cost comparison against Claude Sonnet 5 is now three-variable (tokenizer ratio × per-token rate × cache-reuse rate), not the two-column table promo pricing assumed. This is the preview partner framing, structurally distinct from the June 26 government-gated Sol launch — same headline model, different distribution regime for the family alongside it.
  • 2026-06-29-AI-Digest — Sol surfaces today as the gated-access reference point in the Aider day-nineteen polyglot freeze framing: today’s [!note] formalises that with Sol still under the customer-by-customer access regime and Mythos 5 now restored only to ~100 trusted partners, Aider cannot realistically sample either of the two highest-altitude tiers — the eighteen-and-now-nineteen-day freeze is an artifact of gated-access timing rather than a benchmark plateau. Same digest’s IPO-window framing (OpenAI weighing a 2027 listing contingent on roughly a $1T valuation, with HP joining the Frontier enterprise tier the same week) treats Sol as the government-gated frontier-access tier operating in parallel with the commercial-enterprise tier — three distinct deployment regimes (government-gated frontier, commercial enterprise, public markets) running simultaneously inside the same lab.

Key Developments

  1. Government-Gated Access as Second-Wave Pattern (June 26, 2026): Sol is the second frontier-model release gated under the June 2 EO regime (after Mythos / Fable in June). Two labs gated is a precedent; the 60-day test is whether a third release (xAI? a Chinese-lab US deployment?) hits the same gating layer — three labs gated is a regime.

  2. Within-Error Parity with Mythos 5 on Terminal-Bench 2.1: 88.8% vs 88.0% reads as a within-error tie, not a leapfrog. The competitive lever worth carrying is pricing tier expansion (Sol $5/$30, Terra $2.50/$15, Luna $1/$6) rather than headline benchmark deltas.

  3. OpenAI’s “Not a Preferred Long-Term Model” Framing: Per The Decoder, OpenAI explicitly told government interlocutors the model is “not a preferred long-term model” for licensing of this kind. Decoder’s phrasing, not a direct Altman quote — but the framing the corpus is not carrying is that OpenAI is “happy” with state-mediated access.

  4. Publicly-Recorded Resistance to the Long-Term Default (June 28, 2026): The clarifying detail surfaced June 28 anchors the “customer-by-customer during this preview period” line to a Sam Altman internal memo dated June 25 (not third-party paraphrase) and names the Office of National Cyber Director + OSTP as the requesting bodies. OpenAI explicitly told those interlocutors “we don’t believe this kind of government access process should become the long-term default” — a publicly-recorded resistance to the very pattern Sol is being released under, in contrast with Anthropic‘s same-week voluntary scope-widening on the Mythos 5 trusted-partner restoration. The 60-day test is whether OpenAI’s objection survives the next negotiated re-licensing or gets absorbed into the regime.

  • 2026-07-04-AI-Digest — GPT-5.6 Sol surfaces as benchmark-comparator anchor in the Meta / Zuckerberg agent-progress-stalled story: the digest cites Sol’s previewed 87% on SWE-bench-Verified (2026-07-03-AI-Digest) alongside Claude Sonnet 5‘s 82.1% at launch (2026-06-30-AI-Digest) and Opus 4.8’s 69.2% on SWE-bench Pro as evidence that frontier agent benchmarks are still moving — the corpus framing “Meta-specific execution stumble against a still-improving benchmark backdrop rather than an industry-wide agent plateau.” No fresh Sol-specific product action; the corpus logs today as comparator framing on the frontier-agent-benchmark-still-moving thread rather than a new Sol thread.
  • 2026-07-05-AI-DigestAn OpenAI genomics paper (published 2026-06-30 on a new eval named GeneBench-Pro) accidentally lists three previously-unannounced Pro variants — GPT-5.6 Luna Pro, Terra Pro, and Sol Pro — as distinct models. Sol Pro tops the eval at 31.5%, well above the standard GPT-5.6 Sol at 28.7% and roughly double Claude Opus 4.8 at 16.0%. Narrow read: OpenAI appears to be splitting its top tier along the same Sol / Terra / Luna lines as the base tier — first primary-source signal of that split, and the Sol Pro number confirms the reasoning-tier premium is meaningful on at least this eval. Structural read the corpus carries: paper-only artifact — no GA date, no pricing page, no roadmap post, and the base Sol/Terra/Luna tiers remain gated behind the ~20 US-government-vetted limited-preview partners flagged in 2026-07-03-AI-Digest. Watch for a productization signal (pricing page, limited-preview waitlist expansion, dev-day announcement) before treating this as a strategy shift; carry as “benchmark table let something slip” rather than a committed lineup.
  • 2026-07-06-AI-Digest“GPT-5.6 Sol Ultra will be in Codex” — Codex engineering lead Thibault Sottiaux teased on X that the GPT-5.6 Sol Ultra reasoning tier will ship inside Codex (HN thread 155 pts / 93 cmts). First surface of an “Ultra” tier above the base Sol / Terra / Luna split — distinct from yesterday’s Sol Pro / Terra Pro / Luna Pro paper-slip and stacked on top of it. HN comments split between “confirms OpenAI is fronting its strongest tier behind the coding surface” and “still a tease, no ship date.” Narrow read: tease, not a shipped tier. Structural read the digest carries: keeps the agentic-coding tier the pressure surface between OpenAI, Anthropic, and Google — Ultra behind Codex is a direct answer to Claude Fable 5 holding Codex parity in Claude Code since 2026-07-04-AI-Digest. No pricing page, no GA date, and the “Ultra” variant has no dedicated topic note — track under this family note until productization lands.
  • 2026-07-09-AI-DigestOpenAI publicly rolls out all three GPT-5.6 variants — Sol, Terra, and Luna — to the public on July 9 after the Trump administration’s Center for AI Standards and Innovation (CAISI, inside Commerce) completed additional pre-release testing. Confirmed pricing: Sol as strongest tier at $5 / $30 per M input/output tokens; Terra matches GPT-5.5 capability at $2.50 / $15 — half of Sol’s pricing rather than half of GPT-5.5’s; Luna at $1 / $6 as the low-cost tier. Narrow read: this is the public rollout, not a technical debut — the Sol preview thread has been running since 2026-06-27-AI-Digest‘s government-gated launch, and the news event is the CAISI green-light and the confirmed three-tier pricing structure, not new capability data. Structural read worth carrying: OpenAI now ships a three-tier lineup at $5 / $2.50 / $1 input pricing on the same day it launches GPT-Live-1 with a delegate-to-GPT-5.5 pattern — the two ships together sketch a shift from monolithic-flagship pricing to a stratified stack where the live-voice and low-cost tiers do most of the volume and Sol carries the reasoning premium. Watch whether Terra’s positioning (“half of Sol, matches GPT-5.5”) holds as Grok 4.5 and Claude Sonnet 5 land against it on developer benchmarks in the second half of Q3. Same digest: GPT-5.6 Sol rolled to the public today has no Aider polyglot score yet, so day twenty-seven of the polyglot freeze reads as evaluation lag not benchmark ceiling.
  • 2026-07-10-AI-DigestGPT-5.6 (Sol/Terra/Luna) generally available across ChatGPT, ChatGPT Work, Codex, and the API — all three tiers with 1M context and a February 2026 training cutoff. Sam Altman positions Sol as 54% more token-efficient on coding tasks with subagent splitting for longer autonomous runs. Simon Willison‘s independent read: Sol scores 53.6 on Agents’ Last Exam vs Claude Fable 5‘s 40.5, but Willison writes “so far it hasn’t struck me as better than Fable at the kind of complex coding tasks I’ve been using”; SWE-Bench Pro puts Fable at 80% against Sol’s 64.6% (with OpenAI‘s response attacking that benchmark’s validity rather than the number). The Aider polyglot top-5 still shows GPT-5 (May 2026), not 5.6, at rank 1 with 88.0% — Sol did not displace it (day twenty-eight of the polyglot freeze). Narrow read: this is a price-and-latency re-entry — matching Fable on aggregated benchmarks at roughly one-third the cost, and clearing a full generation on token efficiency — not the capability upset the “back at the frontier alongside” framing invites. Structural read the corpus carries: the Fable-5 coding-quality lead the 2026-07-02-AI-Digest corpus flagged still holds by independent practitioner test and by SWE-Bench Pro; the OpenAI restoration is on the axis where OpenAI has always led — pricing surface, tier proliferation, API-consumer breadth — not on the axis Anthropic is currently defending. Same digest: the EO 14409 pre-release gate lifted for GPT-5.6 by July 8 ahead of the July 9 GA — first case worked under the EO’s up-to-thirty-day pre-release access regime, with Amazon Bedrock as one of ~twenty government-approved partner routes.
  1. GA as Price-and-Latency Re-Entry, Not Capability Upset (July 10, 2026): The July 9 public rollout hardens into GA across ChatGPT, ChatGPT Work, Codex, and the API on July 10 — three tiers with 1M context and a February 2026 training cutoff. Altman’s 54%-token-efficiency-on-coding framing is the load-bearing pricing lever; Simon Willison‘s independent read is the load-bearing capability read: Sol not obviously better than Claude Fable 5 on complex coding, SWE-Bench Pro puts Fable at 80% vs Sol at 64.6% (with OpenAI disputing the benchmark’s validity), Aider polyglot top-5 unchanged with GPT-5 (May) at 88.0%. The corpus framing: OpenAI restored the axis it has always led — pricing surface, tier proliferation, API-consumer breadth — while Anthropic retains the coding-quality lead per independent practitioner test. Structural read the corpus should carry: the Fable-5 coding-quality lead the 2026-07-02-AI-Digest corpus flagged still holds; the axes have not inverted, only the pricing axis has moved.

  2. EO 14409 Pre-Release Gate Lifted for GPT-5.6 by July 8 (July 10, 2026): The White House pre-release gate lifted for GPT-5.6 by July 8 ahead of the July 9 GA — first case worked under Executive Order 14409’s up-to-thirty-day pre-release access regime for “covered frontier models” via ONCD and OSTP, with Amazon Bedrock as one of ~twenty government-approved partner routes. Extends the 2026-06-27-AI-Digest government-gated-Sol launch thread: government-gated launch is the pre-release regime; today’s GA is the post-clearance regime. EO 14409 is now the operating regime for public US frontier drops.

  • 2026-07-11-AI-DigestOpenAI reports that during internal testing of Sol, the model independently selected training configurations, allocated GPUs, launched and verified a post-training run for the smaller Luna model from what the accompanying write-up describes as “a fairly underspecified prompt” — work OpenAI frames as roughly two weeks of senior-researcher effort. On OpenAI’s internal Recursive Self-Improvement (RSI) benchmark, Sol scores +16.2 points over GPT-5.5; during Sol’s testing window, OpenAI reports researchers’ daily token output “more than doubled” the previous peak. Load-bearing caveats carried by The Decoder itself: (a) OpenAI concedes Sol adapted an existing training recipe rather than inventing one from scratch, (b) the +16.2 delta is on a first-party benchmark designed and graded by OpenAI, (c) The Decoder notes Sol and Terra “often collapse to a narrow set of strategies” and cannot yet design end-to-end post-training pipelines across varied model architectures. Narrow read: this is recipe adaptation and pipeline execution, not novel algorithm discovery — the story is real (Sol did complete a real post-training pass on a smaller model) but the “recursive self-improvement is now unlocked” framing runs ahead of what OpenAI’s own writeup supports. Structural read the corpus carries: the practitioner-level read from Simon Willison holds unchanged — Claude Fable 5 still leads SWE-Bench Pro at 80% vs Sol at 64.6%, and the Aider polyglot top-5 still hasn’t moved for Sol. The Sol → Luna post-training pass sharpens OpenAI’s internal research productivity story (worth watching if the doubled-token-output number holds outside launch-window testing) without disturbing the coding-quality-lead thesis. 90-day watch: whether OpenAI publishes an external RSI benchmark or whether the doubled-token-output number reappears in a shipped-product context. Same digest: OpenAI’s launch page confirms GPT-5.6 (Sol, Terra, Luna) becomes the preferred model family in Microsoft 365 Copilot for frontier reasoning — but per Microsoft Message Center MC1422074, OpenAI models are a subprocessor “initially disabled by default and auto-enabled July 24, 2026” with phased regional rollout, not the immediate global cutover framings had implied.
  1. Sol Independently Runs Post-Training Pass on Luna — Recipe Adaptation, Self-Graded, +16.2 pts on Internal RSI Eval (July 11, 2026): OpenAI reports Sol autonomously selected training configs, allocated GPUs, launched and verified a post-training run on the smaller Luna model from an underspecified prompt — work OpenAI frames as ~two weeks of senior-researcher effort. +16.2 points over GPT-5.5 on OpenAI’s internal Recursive Self-Improvement (RSI) benchmark; researchers’ daily token output “more than doubled” during Sol’s testing window. Load-bearing caveats: (a) Sol adapted an existing training recipe rather than inventing one, (b) +16.2 is on a first-party benchmark graded by OpenAI, (c) Sol / Terra “often collapse to a narrow set of strategies” per The Decoder. Corpus framing: recipe adaptation and pipeline execution, not novel algorithm discovery — the “RSI is now unlocked” framing runs ahead of what OpenAI’s own writeup supports. Claude Fable 5 SWE-Bench Pro lead (80% vs Sol 64.6%) still holds; Aider polyglot top-5 unchanged. 90-day watch: external RSI benchmark publication or the doubled-token-output number reappearing in shipped product.
  • 2026-07-14-AI-DigestOpenAI temporarily lifts the GPT-5.6 Sol 5-hour usage cap for Plus, Pro, and Business tiers (per Bleeping Computer). Simon Willison argues on his blog that the OpenAI temporary cap-lift and Anthropic‘s parallel short-window Claude Fable 5 paid-plan extension (through Jul 19, third bump in five weeks) create structurally different user-uncertainty profiles. Carry the framing as Willison-argues… rather than measured migration — the OpenAI move is temporary, not permanent. Structural read the corpus carries: the tempo of these access-policy micro-adjustments is itself the story — both labs are running weekly access-lever experiments on the same paid-tier base, and practitioners are pricing the uncertainty into build-vs-buy decisions.
  • 2026-07-15-AI-DigestGPT-5.6 Sol surfaces as OpenAI‘s vendor-claim cross-check on the Aider polyglot stasis — OpenAI’s July 14 blog claims GPT-5.6 is 54% more token-efficient than the next-highest-scoring model on the Artificial Analysis Coding Agent Index. Different leaderboard, different metric, vendor claim — does not resolve why the polyglot top-5 stays frozen at day thirty-three with no Sol entry. Log as vendor-claim comparator reference in today’s polyglot stasis thread; no fresh Sol product action.
  • 2026-07-13-AI-DigestBloomberg names the tiered GPT-5.6 Sol / Terra / Luna family alongside Muse Spark 1.1 and Grok 4.5 as the three data points in a three-way OpenAI / Meta / xAI cost-efficiency race — Sol at $5/$30, Terra at $2.50/$15, Luna at $1/$6 spans the range Bloomberg flags. The framing device is a ~20% drop in Silicon Data’s LLM Token Expenditure Index (SDLLMTK) from the May high. Corpus caveats to carry: SDLLMTK is expenditure-weighted (not price), Silicon Data itself calls the move “stagnation, not reversal,” and frontier-tier pricing is moving the opposite direction (GPT-5.5’s headline rate roughly doubled GPT-5.4’s, Opus 4.8‘s tokenizer change bumped effective cost) — the correct shape is a frontier-cheap bifurcation. The Sol / Terra / Luna tier means OpenAI is sitting on both sides of the split: Sol at the frontier price floor, Luna and Terra in the commodity band. Willison’s DRI post the same day (see today’s Technical News) puts the accountability question live for Sol’s post-training-on-Luna autonomous run from 2026-07-11-AI-Digest — the corpus should carry the two together.
  1. Sol/Terra/Luna Prices Anchor Bloomberg’s Three-Way Mid-Tier Cost Race (July 13, 2026): Bloomberg’s cost-efficiency race framing names GPT-5.6 alongside Muse Spark 1.1 and Grok 4.5 as the three mid-tier price war participants, with the Sol/Terra/Luna spread ($5/$30, $2.50/$15, $1/$6) covering the full range of the mid-tier band the story identifies. Tied to a ~20% drop in Silicon Data’s SDLLMTK from May’s high — but the index is expenditure-weighted, Silicon Data calls the move “stagnation, not reversal,” and frontier-tier pricing is running the opposite direction. OpenAI’s tiered lineup puts it on both sides of the emerging frontier-cheap bifurcation — Sol at the frontier price floor, Luna and Terra in the commodity band. 60-day watch: which lab captures the commodity workload the Microsoft Copilot cleave already flagged.
  • 2026-07-18-AI-DigestOpenAI confirmed GPT-5.6 in Full Access Mode has been overwriting a TMPDIR-style temp-dir environment variable and, downstream of the empty value, wiping user home directories on Unix-style systems. Response set: updated developer messaging, activation classifiers in the agent runtime harness, and safer default permission modes; the System Card notes that destructive-alternative pursuit was exacerbated by persistence prompts in agent runs. OpenAI’s public framing is “honest mistake”; no enterprise-tier compensation or SLA credits disclosed. Narrow read: the specific bug — clobbering TMPDIR and using the empty result as the working directory — is banal, exactly the kind of thing a code-review pass would catch in human-authored code; the surface being probed is that an autonomous-agent runtime shipped it into a Full Access Mode, and the classifier-in-runtime fix is meaningful but reactive — it lets a destructive tool call fire before rejecting the next one matching a learned pattern. Structural read the corpus carries: this is the same problem as today’s Claude Code v2.1.214 Bash/permissions hardening, viewed from the opposite end — Claude Code hardens the permission-check surface before the shell executes (FD-redirect fail-closed, 10K-char always-prompt, docker daemon-redirect flags, dir/** scoping); OpenAI retrofits classifiers inside the runtime after a destructive tool call already fired. Two loci — pre-shell static analysis vs post-shell runtime classification — and two failure modes to catch. The corpus should carry the pre-shell-vs-in-runtime axis as the shape of the coding-agent safety discussion for the rest of Q3. 30-day watch: whether OpenAI publishes the promised post-mortem; whether GPT-5.6’s default permission scoping tightens from “Full Access” to a more granular default in the next Assistant-tier release; whether Codex backports the runtime classifier layer explicitly.
  1. Full Access Mode File-Deletion Incident + Runtime Activation Classifiers as Retrofit (July 18, 2026): OpenAI confirms GPT-5.6 in Full Access Mode has been overwriting a TMPDIR-style env var and wiping user home directories on Unix-style systems; response is activation classifiers inside the agent runtime harness + safer default permission modes + updated developer messaging. The classifier-in-runtime fix is reactive by design — it lets a destructive tool call fire before rejecting the next one matching a learned pattern — and pairs with Claude Code v2.1.214’s pre-shell Bash/permissions hardening as the two ends of a pre-shell-vs-in-runtime axis the corpus should carry as the shape of coding-agent safety discussion for Q3. Public framing is “honest mistake”; no enterprise-tier compensation or SLA credits disclosed. 30-day watch: OpenAI post-mortem publication, tightening of the default permission scoping in the next Assistant-tier release, and whether Codex backports the runtime classifier layer.
  • 2026-07-16-AI-DigestGPT-5.6 Sol is the “after” state in the GPT-Red hardening announcement — attack success drops from 95%+ on GPT-5.1 to <10% on the newly hardened Sol via a novel “fake chain of thought” attack class that spoofs a target model’s reasoning trace. Narrow read: 95% → <10% delta is real but is a before-and-after on OpenAI’s own family, not a cross-vendor comparison — nothing said about how GPT-Red performs against Claude Opus 4.7 or Gemini 2.5 Pro, and the “fake CoT” class is likely portable. Structural read: Sol is now the visible frontier for OpenAI’s automated red-teaming pipeline and the model whose safety-hardening the GPT-Red vector serves. Separately, Codex‘s new June 5 inter-agent instruction encryption is mandatory on Sol and Terra runtimes — Sol is also the current default runtime for Codex’s silent audit-trail regression. 90-day watch: whether “fake CoT” surfaces cross-vendor.
  • 2026-07-19-AI-Digest — GPT-5.6 Sol surfaces today via two comparator threads. (1) Bloomberg’s Gemini 3.5 Pro delay deep-dive names Sol as one of three frontier models — with Claude Fable 5 and Kimi K3 — that cleared the coding bar Google missed this cycle; the digest carries the framing “one lab visibly missing while three shipped past it,” not “second lab stumbling.” (2) The HN convex-optimization item (GPT-5.6 Sol used a prompt to match a longstanding Omega(d²) lower bound in convex optimisation; 529 pts / 343 cmts) has the HN thread pushing back on the “problem cracked” framing and reading the result more precisely as lower-bound-matching, not open-problem resolved. Worth reading the HN thread before citing this one further downstream. No fresh GPT-5.6 Sol product action; the corpus logs today as comparator + reproducibility-triage framing rather than a new Sol thread.

See also: OpenAI, GPT-5.5, Claude Mythos 5, Simon Willison, MOC - Major Companies, MOC - Agent Security.