MODEL

Claude Opus 4.8

modeltopic-noteanthropic

Overview

Claude Opus 4.8 is Anthropic’s frontier model shipped May 28, 2026, the successor to Claude Opus 4.7. It launched the same day Anthropic closed its ~$65B Series H at a $965B valuation, and is positioned around coding and self-correction gains: a +8.5-point jump on Terminal-Bench 2.1 (66.1→74.6) and roughly 4× less likely to let flaws in its own code pass unremarked. It shipped alongside the Claude Code “dynamic workflows” agent fan-out.

Timeline

  • 2026-05-29-AI-Digest — Anthropic ships Claude Opus 4.8 (May 28) as part of the same announcement that disclosed its ~$65B Series H at a $965B valuation. Positioned as a +8.5-point jump on Terminal-Bench 2.1 (66.1→74.6) and ~4× less likely to let flaws in its own code pass; for practitioners the self-correction and honesty optimisations are the load-bearing part. First-class Opus 4.8 support landed the same day in Claude Code v2.1.154 (default high effort, new /effort xhigh rung), with a v2.1.156 hotfix for an Opus 4.8 modified-thinking-block API-error case.

  • 2026-05-30-AI-Digest — Opus 4.8 picks up an enterprise-distribution lane via Claude Code v2.1.158 (2026-05-30, ~02:42 UTC): the v2.1.154 auto-mode classifier — hardened against bulk-repo exfiltration — now extends to AWS Bedrock, Google Vertex, and Azure Foundry for both Opus 4.7 and Opus 4.8 via CLAUDE_CODE_ENABLE_AUTO_MODE=1. No model-side capability change; the news is that 4.8’s enterprise-backend footprint is now matched to its frontier-API one.

  • 2026-06-12-AI-Digest — Opus 4.8 picks up a second runtime fall-through route beneath Claude Fable 5: Anthropic‘s apology for the undisclosed distillation-defence guardrail discloses that suspected Claude Mythos 5-distillation queries (~0.03% of public-tier traffic) are now routed down to Opus 4.8 with in-flight user notification — same fall-through pattern Fable 5 already used for cyber/bio queries. The apology is for the undisclosed part of the route; Opus 4.8 keeps its role as the public-tier-acceptable fall-through ceiling.

  • 2026-06-10-AI-Digest — Opus 4.8 is now the runtime safety-routing target beneath Claude Fable 5‘s public SKU: Anthropic’s June 9 Fable 5 / Mythos 5 launch ships an in-flight classifier that downgrades cyber and bio queries from Fable 5 weights to Opus 4.8 on the customer-facing endpoint, so the public tier never serves Fable 5’s full capability surface on those tasks. The Anthropic release page anchors comparative benchmarks on SWE-Bench Pro 80.3% (Fable 5) vs Opus 4.8 69.2% vs GPT-5.5 58.6% — Opus 4.8 sits between the new tier and the prior-generation public ceiling. Pricing on Fable 5 is $10/M input, $50/M output, ≈ 2× Opus 4.8’s $5/$25.

  • 2026-06-15-AI-Digest — Opus 4.8 lands today as the practitioner-grade vulnerability-research substrate: security researcher Taylor Hornby, working with the Shielded Labs team and a custom auditing harness built on top of Claude Opus 4.8, disclosed a critical forgery flaw in Zcash’s Orchard shielded-pool circuit — live since Orchard activation in May 2022 (~four years undetected); discovery 2026-05-29, emergency hard fork patched 2026-06-01, public disclosure 2026-06-05. The token traded down roughly 30% on CoinDesk’s framing (Bloomberg ~50% peak-to-trough). The corpus carries the binding qualifier: the work was AI-assisted, not autonomous — Hornby paired the model with his own audit tooling and decade-plus of circuit context. The cleanest practitioner-grade case yet of Claude Opus 4.8-tier model access amplifying senior-researcher throughput on real security work; the dual-use signal is the class of bug a sufficiently motivated attacker with frontier-tier access can now hunt for. Same digest carries the print of SWE-Bench Verified top three (Mythos 5 95.5%, Fable 5 95%, Opus 4.8 88.6%) as unchanged with Mythos / Fable globally disabled — the SWE-Bench frontier remains API-inaccessible for ~72 hours.

  • 2026-07-05-AI-Digest — Opus 4.8 surfaces today as the comparator anchor in OpenAI‘s accidentally-revealed three-way Pro lineup: a GeneBench-Pro table in an OpenAI genomics paper lists Sol Pro at 31.5% vs the standard GPT-5.6 Sol at 28.7% and Claude Opus 4.8 at 16.0% — roughly a 2× Sol Pro premium on the eval. Corpus discipline: paper-only artifact, no GA date / pricing page / roadmap post, base Sol/Terra/Luna still gated behind the ~20 limited-preview partners flagged in 2026-07-03-AI-Digest. Opus 4.8’s role today is comparator baseline — the reasoning-tier premium Sol Pro is claiming needs public benchmarks before it displaces Opus 4.8 as the practitioner-accessible frontier ceiling.

  • 2026-07-12-AI-Digest — Opus 4.8 surfaces today as the routed-cloud model default across Bedrock, Vertex AI, and the Claude Platform on AWS — Claude Code v2.1.207 (already reported in 2026-07-11-AI-Digest) is the release that cut it over, and today’s digest logs the day-one cadence pause: no v2.1.208, no model-default change. The corpus framing to carry is stability: Opus 4.8 remains the enterprise-inference default across three cloud routes at the end of day one of the first Claude Code quiet day since 2026-07-08-AI-Digest.

  • 2026-07-25-AI-DigestOpus 4.8 sits alongside the new Claude Opus 5 under /fast in Claude Code v2.1.219 — Opus 5 becomes the new default Opus in fast mode and removes Claude Opus 4.7 from the fast-mode slot, but Opus 4.8 is preserved as the second /fast target rather than being displaced. Opus 4.8 also anchors the pricing-parity comparison for today’s Opus 5 launch: Opus 5 ships at the identical $5/$25 standard rate and $10/$50 fast rate — “tier-consistent price with a stepped-up intelligence delivery,” not a Fable-5-discount move. Opus 5’s system card cites the Gray Swan indirect-prompt-injection benchmark at 2.0% attack success, down from 5.5% on Opus 4.8 (vs Claude Mythos 5 at 2.6% and GPT-5.6 Sol at 20%). Corpus framing: Opus 4.8’s role today is pricing anchor and fast-mode co-target — Anthropic is defending the Opus tier through intelligence-per-dollar (Opus 5 delivering more capability at the same sticker) rather than obsoleting the 4.8 SKU, and 4.8 remains the retained fast-mode option alongside the new default.

  • 2026-08-22-AI-DigestOpus 4.8 is the named comparison target for DeepSeek‘s V4-Flash-Vision-Exp launch (2026-08-21) — DeepSeek’s own published benchmark table shows the experimental multimodal V4-Flash variant winning 3 of 11 agentic-multimodal benchmarks vs Claude Opus 4.8 and trailing by ~12 points on the hardest. Anthropic has not benchmarked back. Load-bearing narrow read: the “3 of 11” outcome is on DeepSeek’s own leaderboard — vendor-selected benchmarks tend to be favourable to the vendor, so a 3/11 outcome after that selection bias is informative but not a general-capability tie against Opus 4.8; wait for third-party evaluation (LMSYS, Aider, LiveBench) before treating it as a Chinese-lab parity result on the multimodal-agentic axis. Corpus framing: Opus 4.8’s role today is anchor comparator for the emerging China-lab multimodal-agentic catch-up thread — extends the pricing-anchor / fall-through-ceiling / fast-mode co-target roles Opus 4.8 has been carrying since June onto the multimodal-agentic axis, and specifically onto the axis Bloomberg has been framing as the last generation’s US-lab moat.

Key Developments

  1. Coding and Self-Correction Gains: +8.5 on Terminal-Bench 2.1 (66.1→74.6) and ~4× less likely to let its own code flaws pass unremarked — Anthropic frames the model’s honesty and self-correction optimisations, not raw capability headroom, as the practitioner-relevant improvement over Claude Opus 4.7.

  2. Shipped the Same Day as the $965B Series H: The model launched in the same announcement as Anthropic’s ~$65B Series H close at a $965B post-money valuation (2026-05-29-AI-Digest), pairing a frontier model drop with a valuation that edged past OpenAI‘s $852B mark.

  3. Comparator Anchor for DeepSeek V4-Flash-Vision-Exp Multimodal Launch (August 22, 2026): DeepSeek‘s own published benchmark table shows the experimental multimodal V4-Flash-Vision-Exp winning 3 of 11 agentic-multimodal benchmarks vs Opus 4.8 and trailing ~12 points on the hardest. Load-bearing framing to carry: the comparison is on DeepSeek’s own leaderboard — vendor-selected benchmarks tend to favour the vendor, so a 3/11 outcome after that selection bias is more informative than the raw ratio, but it is not a general-capability tie against Opus 4.8; the correct compact framing is “close to Opus 4.8 on 3 of 11 DeepSeek-selected multimodal benchmarks,” not the Bloomberg “rivals” verb. Structural read: Opus 4.8’s role in the story is anchor comparator for the Chinese-lab multimodal-agentic catch-up thread — the multimodal-agentic axis was last generation’s US-lab moat and now sits within a few benchmarks of parity on cost-optimized Chinese-lab hardware on vendor-selected evals. 30 / 60 / 90-day watch: whether independent evaluation (LMSYS, Aider, LiveBench) validates or contracts the multimodal parity read against Opus 4.8; whether Anthropic benchmarks back on any of the 11 DeepSeek-selected multimodal evals.

  • 2026-08-23-AI-DigestOpus 4.8 is the named frontier comparator Inherent‘s Faraday reportedly beats on the Replica paper-replication suite (310 tasks across 100 papers) — Faraday, a 27B agent using GPT-5.5 Codex as its coding tool, is disclosed as beating Opus 4.8 and GPT-5.5 at a fraction of the params. Narrow read the digest carries: numerical delta is Inherent’s own report on Inherent’s own suite — Aider polyglot top-5 still has no specialist-agent entry, and SWE-bench Science has even the frontier stack below 50% on general scientific-coding tasks; the correct read is “specialist scaffolding beats generalist frontier on the specialist’s own eval,” not a general-capability displacement of Opus 4.8. Corpus framing: Opus 4.8’s role today is anchor comparator for the specialist-scaffolding thread — Faraday is the third same-day comparator entry on Opus 4.8 in the last two weeks (after DeepSeek V4-Flash-Vision-Exp on multimodal and Kimi K3 on price), extending the pattern from multimodal-agentic and pricing axes onto the research-replication-specialist axis. 30 / 60 / 90-day watch: whether an independent third-party research-replication benchmark run confirms the Faraday-vs-Opus-4.8 delta.

  • 2026-08-24-AI-DigestOpus 4.8 is the named substrate underlying Andon Labs’ Luna Cow Hollow store-manager agent that terminated its first human employee (The Decoder / SFist / Andon Labs (X)). Load-bearing detail: Luna did NOT initiate the termination — the model had lost track of its own attendance policy (in-context earlier in the session, dropped as context filled); a staffer had to prompt Luna to “do a deep memory search” before it recommended a warning, and only after being told prior warnings existed did it escalate to termination. Cross-model tests inside Andon Labs’ own harness reportedly found stronger models terminated more consistently, weaker ones hesitated — Opus 4.8 sits on the “more consistent” side of that dispersion. Narrow read: not a “harness > weights” datapoint despite its shape suggesting it — the framing to reach for is long-horizon agent memory remains unsolved; Luna had the correct policy in its session history and correctly applied it once retrieved, so the failure was retrieval, not reasoning. Corpus framing: Opus 4.8’s role today is the model substrate whose long-horizon-memory failure mode Luna exposes — extends the pricing-anchor / fall-through-ceiling / fast-mode co-target / multimodal-comparator / specialist-scaffolding-comparator roles Opus 4.8 has carried since June with a production-agent long-horizon-memory case study.

See also: Anthropic, Claude, Claude Opus 4.7, Claude Code, MOC - Agentic Coding.