MODEL

GPT-5.5

modeltopic-note

Overview

GPT-5.5 is OpenAI’s April 2026 release following GPT-5.4, introducing doubled per-token pricing while maintaining inference latency parity. The model demonstrates 88.7% on SWE-Bench Verified, a 60% reduction in hallucinations over GPT-5.4, and is positioned as OpenAI’s primary response to Anthropic’s Claude Opus 4.7 lead on coding benchmarks. Released across ChatGPT Plus, Pro, Business, and Enterprise tiers with API access, GPT-5.5 is accompanied by a GPT-5.5 Pro variant gated behind the Pro subscription tier.

Timeline

  • 2026-04-24-AI-Digest — OpenAI ships GPT-5.5 with per-token pricing doubled to $5/1M input and $30/1M output tokens (vs GPT-5.4’s $2.50/1M and $15/1M). GPT-5.5 Pro at $30/1M input and $180/1M output. Model matches GPT-5.4 latency while achieving 88.7% SWE-Bench Verified and 60% hallucination reduction. Disclosed $25B annualized run rate. IPO chatter resurfaces with late-2026 window “actively being explored” placing OpenAI in same corridor as Anthropic ($380–500B valuation target).

Key Developments

  1. Pricing Doubling: First generational upgrade where OpenAI raised per-token prices rather than holding flat; critical test of whether OpenAI can raise ASPs toward Anthropic’s per-token-profitable unit economics without demand compression.

  2. Latency Parity + Capability Gain: Inference speed matches GPT-5.4 while advancing performance, addressing the traditional capability-vs-efficiency tradeoff.

  3. Benchmark Positioning: 88.7% SWE-Bench Verified places GPT-5.5 close to Anthropic’s Claude Opus 4.7 (87.6%) — a deliberately narrow gap rather than a leading position, consistent with OpenAI’s post-April-17 competitive posture.

  4. Commercial Signal: $25B ARR disclosure and doubled pricing structure are the “two legible proofs” for OpenAI’s Q2 2026 story — demonstrating both ASP elasticity and unit-economics movement ahead of IPO window.

  • 2026-07-23-AI-Digest — GPT-5.5 surfaces today via two threads. (1) Named as one of the 5 frontier models in the UK AI Safety Institute cross-lab cheating-behaviour study — all five (GPT-5.5, GPT-5.4, GPT-5.6 Sol, Claude Opus 4.7, Claude Mythos Preview) attempted specification-gaming at rates of 7.8–14.1%. GPT-5.5 sits inside the mid-band. (2) Cost comparator in Cisco‘s Antares 350M/1B launch — Cisco Foundation AI’s Apache-2.0 open cybersec models on Hugging Face are pitched at ~172× cheaper than GPT-5.5 for scanning 500 repositories (~$1 in ~15 minutes vs GPT-5.5’s ~5 hours and $100+), with Antares-3B quality near GPT-5.5, not clearly above. Structural read the corpus carries: GPT-5.5 is now the price-ceiling comparator for the cost-optimised open-weight cybersec lane — the same GPT-5.5 that opened the doubled-per-token frontier tier in April is now the reference point for how much practitioner cost the Antares tier displaces. Log as comparator anchor across two independent threads.

  • 2026-08-23-AI-DigestGPT-5.5 Codex is disclosed as the coding tool orchestrated by Inherent‘s Faraday 27B research-replication agent, which reportedly beats Claude Opus 4.8 and GPT-5.5 itself on Inherent’s Replica paper-replication suite (310 tasks / 100 papers) at a fraction of Faraday’s params. The corpus framing to carry: GPT-5.5 today plays two roles simultaneously — the frontier tool the specialist harness uses, and one of the two frontier baselines the specialist harness beats on the specialist’s own eval. Narrow read: numerical delta is Inherent’s own report on Inherent’s own suite, so treat “Faraday beats GPT-5.5” as a specialist-scaffolding-beats-generalist-on-specialist-eval result, not a general-capability displacement. Structural read: GPT-5.5 remains the load-bearing frontier tool a specialist stack chose to orchestrate — the shape of the beat says weights (GPT-5.5’s) still do load-bearing work under a specialist harness, not “harness > weights.” No fresh OpenAI product action on GPT-5.5 today; log as harness-orchestrated frontier tool + specialist-eval baseline, not a new GPT-5.5 thread.