MODEL

DeepSeek V4 Pro

modeltopic-note

Overview

DeepSeek V4 Pro is a frontier reasoning model from DeepSeek, a Chinese AI lab. In May 2026, DeepSeek V4 Pro demonstrated strong cost-efficiency on agentic benchmarks, positioning at a fraction of GPT-5.2’s pricing while achieving comparable performance on specific tasks.

Timeline

  • 2026-05-06-AI-Digest — A Resources flair post on r/LocalLLaMA reported that DeepSeek V4 Pro ties GPT-5.2 on FoodTruck Bench — a 30-day agentic benchmark with persistent memory — at roughly 17× lower cost per million tokens ($0.435 / $0.87 input/output for V4 Pro versus $1.75 / $14 for GPT-5.2). The headline framing in the thread is that “the China–US frontier gap has compressed to ten weeks.” Treat that framing carefully: independent evals put DeepSeek’s broader capability lag closer to 6–8 months on reasoning and 12+ months on multimodal and code, so what compressed is one specific benchmark profile (long-horizon agentic with memory), not the overall capability surface. The 17× cost ratio is the load-bearing number for AI-startup unit economics — at GPT-5.2 prices, agentic loops with persistent memory burn through margins fast; at V4 Pro prices the same loop becomes price-insensitive. The capability-parity claim is benchmark-specific and shouldn’t be over-extrapolated; the cost story is structural.

  • 2026-05-11-AI-Digest — r/LocalLLaMA post (“I have DeepSeek V4 Pro at home”, 245 upvotes, 122 comments) documents a successful Q4_K_M run on an EPYC 9374F workstation (12×96 GB RAM, single RTX PRO 6000 Max-Q) using a community CUDA fork of llama.cpp with modified Q4_K_M support. Worked out of the box. The frontier-class MoE model in this weight class is now self-hostable on prosumer hardware budgets — continues narrowing the “you need a cluster for this” envelope.

  • 2026-05-24-AI-DigestDeepSeek formalises the 75% promotional discount as the permanent list rate: $0.435/M input (cache miss), $0.003625/M (cache hit), $0.87/M output. The previous list prices ($1.74/M input, $3.48/M output) are retired; the long-running promo becomes the standard rate. Against GPT-5.5‘s $5/M input and $30/M output that’s roughly 11.5× cheaper on input and 34× cheaper on output, with the cache-hit input rate at sub-cent-per-million economics no US frontier lab is publishing. The price formalisation locks in the China-vs-US frontier-API gap at the ~10–35× range rather than the 3–5× US analysts had assumed would re-converge once promo pricing ended.

  • 2026-05-25-AI-Digest — V4-Pro becomes the load-bearing economics for Reasonix, a community / third-party terminal coding agent (esengine GitHub org, MIT-licensed, npm reasonix, ~5.5k★) engineered explicitly around V4-Pro’s prefix-cache behaviour. Reasonix claims a 99.82% prefix-cache-hit rate and ~93% cost savings against Claude Code equivalents — direct validation that the cache-hit input rate ($0.003625/M) is the price point that changes the architecture of how coding agents structure their context. Lands the day after permanent-pricing went live, on the HN front page (495 pts / 208 cmts). The read is demand-side: practitioners built a working agent around V4-Pro’s prefix cache the day after the pricing formalised, not that DeepSeek is shipping a first-party agent.

  • 2026-06-08-AI-Digest — HN thread “DeepSeek V4 Pro beats GPT-5.5 Pro on precision” (167 pts / 48 cmts) cites a runtimewire piece claiming V4 Pro edges out GPT-5.5 Pro on precision-focused benchmarks. The digest hedges hard on the framing: the runtimewire piece isn’t a known publication, the win appears task-specific (one bug-finding eval), GPT-5.5 still leads SWE-bench Pro ~58.6% vs V4 Pro’s 55.4%, and DeepSeek V4 Pro is absent from today’s Aider polyglot top-5. The signal is cost-disruption, not capability parity — pair with DeepSeek topping Ramp’s June trending-vendors index for the procurement read. NIST CAISI still has V4 Pro roughly eight months behind frontier reasoning; the task-specific win does not change that overall picture.

Key Developments

  1. FoodTruck Bench Parity with GPT-5.2: Achieves comparable performance on 30-day agentic benchmark with persistent memory at 17× lower cost.

  2. Cost-Efficiency Leadership: Pricing structure ($0.435 / $0.87 per million tokens) enables price-insensitive agentic loops that become margin-critical economics at GPT-5.2 pricing levels.

  3. Capability-vs-Cost Positioning: Benchmark-specific parity does not represent broader capability parity (6–8 month lag on reasoning, 12+ months on multimodal/code); cost advantage is the structural differentiation for agentic workloads.

  4. Prosumer Home Deployment (May 2026): Q4_K_M run on single-RTX PRO 6000 Max-Q workstation working out of the box establishes that frontier-class MoE models at this weight class are now self-hostable on prosumer hardware budgets, validating the continuing local-vs-frontier curve compression.

  5. Permanent 75% Discount as List Pricing (May 24, 2026): The promo rate becomes the permanent list rate — $0.435/M input cache-miss, $0.003625/M cache-hit, $0.87/M output. Roughly 11.5× cheaper input and 34× cheaper output than GPT-5.5. The signal is structural: the China-vs-US frontier-API price gap has been locked in at the ~10–35× range rather than the 3–5× re-convergence US analysts had assumed.

  • 2026-08-13-AI-DigestDeepSeek shipped a V4 Pro 0813 checkpoint on OpenRouter with no blog post or tweet — the only signal was the API docs update, flagged by Simon Willison. HN thread ran 827 pts / 326 cmts as the day’s heaviest evaluation surface, where the comparison against GPT-5.6 and Grok 4.6 played out in real time. Narrow read to carry: stealth-ship shape — API-docs-only surface with no marketing means the release is being evaluated on the practitioner-benchmark axis rather than the vendor-pitch axis. Structural read: V4 Pro 0813 is one of three frontier-adjacent drops in three days (Grok 4.6 same day, Muse Glimmer on Aug 10) — all pricing or distributing to undercut the Anthropic / OpenAI price bracket rather than beat them on a headline benchmark, with V4 Pro 0813 as the silent-API-docs variant of the release-shape triptych (Grok 4.6 = coordinated multi-surface distribution, Muse Glimmer = Apache 2.0 weights drop). Extends the 2026-05-24-AI-Digest permanent-list-pricing thread and the 2026-08-09-AI-Digest pre-IPO price-hike signal with the stealth-checkpoint-refresh leg — DeepSeek is now visibly refining V4 Pro on the API surface while telegraphing pricing changes on the same product. 30 / 60 / 90-day watch: whether independent benchmark scores for V4 Pro 0813 land inside the HN discussion window; whether DeepSeek publishes a blog post or model card retrospectively; whether the stealth-ship becomes the DeepSeek default cadence given the pre-IPO pricing signal.
  1. V4 Pro 0813 Stealth Ship on OpenRouter — 827-pt HN Thread as Primary Evaluation Surface (August 12, 2026, covered August 13): DeepSeek’s V4 Pro 0813 checkpoint lands on OpenRouter with no blog post or tweet — API docs update was the only signal, Simon Willison surfaced it publicly, HN thread ran 827 pts / 326 cmts. Load-bearing framing to carry: the stealth-ship shape is the story — API-docs-only distribution is provisioning-team notice, not marketing, and the release is being evaluated on practitioner-benchmark axes rather than vendor pitch. Structural read: V4 Pro 0813 is the silent-API-docs variant in a three-drop / three-day frontier-undercut cluster (Grok 4.6 = coordinated multi-surface distribution; Muse Glimmer Aug 10 = Apache 2.0 weights drop). All three drops target price/distribution rather than headline capability — the competitive front is moving from which model is best to which model is cheap enough to route the median agent call to. 30 / 60 / 90-day watch: whether independent benchmarks for V4 Pro 0813 land during the HN discussion window; whether DeepSeek retroactively publishes a blog / model card; whether stealth-ship becomes the default cadence.
  • 2026-08-14-AI-DigestDeepSeek made V4 Pro generally available on the DeepSeek API alongside the DeepSeek Harness v0.1 developer preview at higher per-token rates than V4 (per VentureBeat). No explicit pricing table disclosed in the launch materials the digest cites; the higher-than-V4 direction of travel matches 2026-08-09-AI-Digest‘s pre-IPO API price-hike signal. Narrow read to carry: V4 Pro is now the paired-model anchor to the MIT-licensed DeepSeek Harness reference agent runtime — DeepSeek is shipping (model + harness) as one bundle on the same news day. Structural read the corpus carries: pairs with the 2026-08-13-AI-Digest stealth-ship 0813 checkpoint and the 2026-08-09-AI-Digest pre-IPO pricing signal as three same-week V4 Pro moves — checkpoint refresh, price direction, and now a first-party open agent runtime the model can be evaluated inside. 30 / 60 / 90-day watch: whether DeepSeek publishes the explicit V4-vs-V4-Pro API price table on the pricing page; whether independent practitioner reviews cross-evaluate V4 Pro through DeepSeek Harness against Claude Code / Kitesurf harnesses on the same task suite; whether the (model + harness) bundle emerges as the default DeepSeek go-to-market shape.
  1. V4 Pro Ships on DeepSeek API Alongside MIT-Licensed DeepSeek Harness at Higher Rates Than V4 (August 13, 2026): V4 Pro becomes the paired-model anchor for the DeepSeek Harness v0.1 open agent runtime — a (model + harness) release shape that matches the pattern Kitesurf and Claude Code have set. Per-token rates higher than V4, direction of travel consistent with the 2026-08-09-AI-Digest pre-IPO pricing signal. Structural framing to carry: three same-week V4 Pro moves — 2026-08-13-AI-Digest stealth 0813 checkpoint refresh, pre-IPO price direction, and today’s harness-paired API GA — position V4 Pro as the load-bearing reference DeepSeek is choosing to evaluate its stack around, not the model most likely to top a benchmark. 30 / 60 / 90-day watch: explicit V4-vs-V4-Pro pricing-page disclosure; whether the (model + harness) bundle becomes DeepSeek’s default go-to-market; whether independent practitioners cross-benchmark V4 Pro inside DeepSeek Harness against Claude Code / Kitesurf on the same task suite.
  • 2026-08-17-AI-DigestV4 Pro output tokens climb to $3.96/M peak ($1.98 off-peak) under the DeepSeek repricing that took effect 16:00 UTC 2026-08-16. Part of the broader V4 API rate revision spanning +57% to over +1,100% across token types under the new peak/off-peak split — the load-bearing V4-Pro-specific numbers are the peak $3.96/M output and off-peak $1.98/M output. Bloomberg framing (not DeepSeek’s) is capacity-driven and pre-IPO; DeepSeek has not issued a public statement of intent. Narrow read: the 11× ceiling that shows up in coverage is a different token class’s peak, not V4 Pro output — V4 Pro output specifically moves at a lower multiplier under the new schedule, and the corpus should carry the tier-specific numbers rather than the headline range. Structural read: V4 Pro’s specific price move lands V4 Pro as the frontier-tier-priced member of the V4 family — its output rates are now materially closer to Claude Sonnet 5 and GPT-5.6 Sol tier economics than to earlier V4 Pro pricing, and the tier-differentiation from V4 Flash widens under the new schedule. Extends the 2026-05-24-AI-Digest permanent-list-pricing thread (the $0.435/$0.87 pattern) with the effective-rate step-up leg — V4 Pro’s Aug 2026 pricing floor is no longer the reference cost point it was through May–July. 30 / 60 / 90-day watch: whether DeepSeek publishes an explicit V4-vs-V4-Pro pricing-page update; whether Reasonix-style prefix-cache-optimized agents preserve their cost-savings claim under the new peak/off-peak schedule; whether V4 Pro’s paired-model role with DeepSeek Harness holds through the pricing revision.
  1. V4 Pro Output $3.96/M Peak, $1.98/M Off-Peak Under Effective 2026-08-16 Repricing (August 17, 2026): V4 Pro output climbs materially under the new peak/off-peak split — $3.96/M peak, $1.98/M off-peak. Load-bearing framing to carry: the 11× ceiling in coverage is a different token class’s peak, not V4 Pro output — V4 Pro moves at a lower multiplier under the new schedule; carry the tier-specific numbers rather than the headline range. Structural read: V4 Pro is now the frontier-tier-priced member of the V4 family — output rates materially closer to Claude Sonnet 5 and GPT-5.6 Sol tier economics than earlier V4 Pro pricing, and V4-vs-V4-Flash tier differentiation widens under the new schedule. Extends the 2026-05-24-AI-Digest permanent-list-pricing thread with the effective-rate step-up leg — V4 Pro’s Aug 2026 pricing floor is no longer the reference cost point through May–July. 30 / 60 / 90-day watch: explicit V4-vs-V4-Pro pricing-page update; Reasonix-style prefix-cache-optimized agents’ cost-savings claim under the peak/off-peak schedule; V4 Pro’s paired-model role with DeepSeek Harness through the pricing revision.