MODEL

DeepSeek-V4-Flash

modeltopic-notedeepseekopen-weightsamd

Overview

DeepSeek-V4-Flash is a current-generation frontier open-weights model from DeepSeek, sitting in the V4 family alongside DeepSeek v4 and DeepSeek V4 Pro. Its first substantive surfacing in the AI Digest corpus is via a practitioner-grade port to AMD MI300X — a concrete data point on whether the AMD inference stack is closing the gap on the current frontier open-weights cohort.

Timeline

  • 2026-06-03-AI-Digest — Surfaces via Fergus Finn’s Hacker News write-up (fergusfinn.com, 94 points · 11 comments) on porting DeepSeek-V4-Flash inference to AMD MI300X — including FP8 fnuz vs OCP mismatches, AITER gaps on gfx942, and ROCm helper work needed to get the model serving. Load-bearing for the “CUDA moat” thread: this is one of the cleaner practitioner data points to date on how much friction remains to bring a current frontier open-weights model up on a non-NVIDIA accelerator end-to-end. The write-up is granular enough to be useful as a reference for anyone attempting the same port.

Key Developments

  1. First Practitioner-Grade MI300X Port Write-Up in the Corpus: Fergus Finn’s HN-front-page post is the cleanest single artifact this corpus has on what’s actually required to bring a current frontier DeepSeek model up on AMD inference silicon — concrete pain points (FP8 fnuz vs OCP, AITER gaps, ROCm helpers) rather than benchmark talking points.

  2. Open-Weights Frontier Cohort Member: DeepSeek-V4-Flash sits in the V4 family (DeepSeek v4, DeepSeek V4 Pro) and the Flash positioning suggests an efficiency tier of the open-weights frontier — the right peer set for MI300X port economics is current-generation open-weights frontier models, not the proprietary closed-weights cohort.

  • 2026-07-23-AI-DigestDeepSeek-V4-Flash surfaces as the domain-baseline in the SLAI T-Rex paper (arXiv:2607.20145, ▲25) — full-parameter post-training framework running on Huawei Ascend NPU SuperPOD produces an Operations-Research specialised variant that beats base DeepSeek-V4-Flash by 11.27pp on 71.81% zero-shot Pass@1 and beats GPT-5-family Mini by 3.98pp on the same eval. Narrow read: DeepSeek-V4-Flash is the pre-fine-tuning baseline the paper measures against, not the paper’s headline result. Structural read the corpus carries: the Flash tier of the V4 family is now the load-bearing reference-baseline for non-NVIDIA trillion-scale post-training — the same efficiency-tier positioning that made the 2026-06-03-AI-Digest MI300X port write-up load-bearing is what makes it a useful baseline for Ascend-NPU-based fine-tuning. Extends the open-source-models thread with a frontier-open-baseline datapoint on the Ascend-hardware axis. Log as reference-baseline appearance on a same-week non-NVIDIA post-training result.
  • 2026-08-01-AI-DigestDeepSeek ships V4 Flash 0731 — a 304B-parameter model at $0.14/$0.28 per M tokens ($0.014/M cache-hit) — ranking ahead of MiniMax M3 (428B) on the Artificial Analysis Intelligence Index at 40, sharing that bucket with Thinking Machines Lab‘s same-week Inkling Small. Simon Willison‘s hands-on note is that this is plausibly the cheapest “intelligent” model available per input token but the default reasoning setting is mediocre — high reasoning effort delivers the good outputs but inflates output-token counts (~45K/task at max effort). Narrow read: cheapest-per-input-token, softens on per-completed-task once the reasoning-effort tax is priced in. Undercuts yesterday’s GPT-5.6 Luna cut ($0.20/$1.20) on both sides at headline price. Structural read the corpus carries: small-reasoning-model is now an emerging benchmarking bucket at Artificial Analysis (V4 Flash 0731 at 304B dense vs Inkling Small at 276B / 12B-active is a very different scaling shape at the same Intelligence Index score) — not a defined parameter cutoff yet. 30-day watch: whether a third entrant lands in the small-reasoning / low-price / open-weights bucket that would move this from co-emergence to a genuine category.
  1. V4 Flash 0731 Ships as 304B at $0.14/M Input in the Emerging Small-Reasoning Bucket (August 1, 2026): 304B-parameter model at $0.14/$0.28 per M tokens ($0.014/M cache-hit) ranking ahead of MiniMax M3 (428B) at Artificial Analysis Intelligence Index 40 — same score as Thinking Machines Lab‘s Inkling Small (276B / 12B active) in the same 48-hour window. The disciplined framing to carry: small-reasoning-model is an emerging benchmarking bucket, not two independent announcements, and also not yet a defined size class (304B dense vs 12B-active MoE is a very different scaling shape). Simon Willison‘s hands-on adds the reasoning-effort-tax read: cheapest-per-input-token, softer per-completed-task at max effort (~45K tokens/task output). Undercuts yesterday’s GPT-5.6 Luna $0.20/$1.20 cut on headline price. Extends the 2026-06-03-AI-Digest MI300X-port thread and the 2026-07-23-AI-Digest SLAI T-Rex baseline thread by anchoring the Flash tier as the price-and-benchmark reference point for the emerging small-reasoning bucket.
  • 2026-08-08-AI-DigestDeepSeek V4 Flash 0731 lands on the ARC Prize board and hits HN’s front page (~534 pts / ~318 cmts on arcprize.org/results/deepseek-v4-flash-0731) — the 0731 checkpoint gets its benchmark surfacing on ARC’s public results page with heavy practitioner discussion. Narrow read: the ARC Prize placement is the news, not a fresh model release — the 0731 checkpoint shipped in 2026-08-01-AI-Digest. Structural read the corpus carries: frontier open-weights HN reception cadence — DeepSeek open-weights releases remain a core corpus thread since 2026-08-02-AI-Digest, and ARC Prize placement is one of the few benchmarks not yet frozen by the Aider leaderboard stall. Log as benchmark-surface continuation on the 0731 checkpoint, not a new release.

  • 2026-08-09-AI-DigestV4 Flash surfaces today as the pricing anchor in DeepSeek‘s Aug 6 developer email warning of a substantial cross-service API price increase — the second pricing move inside a month. Current prices remain V4 Flash at $0.14 / $0.28 per M input / output tokens versus Kimi K2.5 at $3 / $15 as the load-bearing frontier-tier comparator. Bloomberg frames the DeepSeek move as pre-IPO commercialization pivot; no specific hike percentage disclosed and no effective date named. Narrow read: V4 Flash is the pre-hike anchor the “just use DeepSeek” default has been priced against, and the impending hike specifically weakens that default for the Flash tier. Structural read: the impending price move on V4 Flash is DeepSeek-specific, not China-wide — Qwen 3.5 Flash still lists at $0.10 / $0.40 per M and Kimi K2.5 sits at $0.60 / $3, and Apidog’s H1 2026 tracking counted six Chinese-lab price cuts in the first half. The compute-economics assumption that weakens today is the “just use V4 Flash” default specifically, not the broader China open-weight low-cost story. 30 / 60 / 90-day watch: whether V4 Flash’s new sticker lands in a formal pricing page update; whether inference-cost-sensitive agentic architectures start migrating provider defaults away from V4 Flash in weekly practitioner posts.

  • 2026-08-17-AI-DigestV4 Flash output tokens moved from $0.28 → $1.32 per million at peak ($0.66 off-peak) under the DeepSeek repricing effective 16:00 UTC 2026-08-16. Part of the broader V4 API rate revision spanning +57% to over +1,100% across token types under the new peak/off-peak split — V4 Flash output specifically climbs ~4.7× at peak, ~2.4× off-peak. Bloomberg framing (not DeepSeek’s) is capacity-driven and pre-IPO. Narrow read: the specific V4 Flash output multiplier is ~4.7× peak, not the 11× ceiling that shows up in headline coverage — carry the tier-specific numbers rather than the headline range. Structural read: the V4 Flash effective sticker predicted by the 2026-08-09-AI-Digest Aug 6 developer-email telegraph has now landed, and the tier-specific ~4.7× peak multiplier is the load-bearing datum practitioners running V4 Flash-anchored inference pipelines need to reprice against — the “just use V4 Flash” default the note has been anchored on is now materially weaker at peak windows, though off-peak remains competitive with Qwen 3.5 Flash at $0.10/$0.40 per M. Extends the 2026-08-09-AI-Digest pre-hike-telegraph leg with the effective-rate-landing leg — the ARC Prize board result from 2026-08-08-AI-Digest on the 0731 checkpoint stands, but the cost-per-completed-task calculus around it is now materially different at peak windows. 30 / 60 / 90-day watch: whether the peak/off-peak split flushes hobbyist and batch workloads off the platform specifically at the Flash tier; whether the Aug 6 developer-email framing gets a follow-up pricing-page update from DeepSeek; whether Qwen 3.8 27B self-hosting emerges as the practitioner-visible Flash-tier substitution.

  1. V4 Flash Output $0.28 → $1.32/M Peak ($0.66 Off-Peak) Under Effective 2026-08-16 Repricing (August 17, 2026): V4 Flash output specifically climbs ~4.7× at peak, ~2.4× off-peak under the new peak/off-peak split (part of the broader +57% to over +1,100% cross-service range). Load-bearing framing to carry: the tier-specific V4 Flash output multiplier is ~4.7× peak, not the 11× ceiling that shows up in headline coverage — hold the tier-specific numbers rather than the headline range. Structural read: the V4 Flash effective sticker predicted by the 2026-08-09-AI-Digest Aug 6 developer-email telegraph has now landed — the “just use V4 Flash” default the note has been anchored on is materially weaker at peak windows, though off-peak remains competitive with Qwen 3.5 Flash at $0.10/$0.40 per M. Extends the 2026-08-09-AI-Digest pre-hike-telegraph leg with the effective-rate-landing leg. 30 / 60 / 90-day watch: peak/off-peak split flushing hobbyist / batch workloads off the Flash tier; formal pricing-page update from DeepSeek; whether Qwen 3.8 27B self-hosting emerges as the practitioner-visible Flash-tier substitution.
  • 2026-08-22-AI-DigestDeepSeek on 2026-08-21 launched V4-Flash-Vision-Exp — an experimental multimodal variant of DeepSeek-V4-Flash that can interpret visual prompts alongside text — live on the DeepSeek API. On DeepSeek’s own published benchmark table, the model wins 3 of 11 agentic-multimodal benchmarks vs Claude Opus 4.8 and trails by ~12 points on the hardest. Anthropic has not benchmarked back. Bloomberg framed the release as another data point in the Chinese-lab catch-up trend it has been reporting all week alongside Moonshot AI and Z.ai‘s coding coverage. Narrow read: do not lift the Bloomberg “rivals” verb — the correct compact framing is “close to Opus 4.8 on 3 of 11 DeepSeek-selected multimodal benchmarks”; vendor-selected benchmarks tend to be favourable to the vendor, so a 3/11 outcome after that selection bias is more informative than the raw ratio suggests but is not a general-capability tie. Structural read the corpus carries: the multimodal-agentic axis — last generation’s US-lab moat — is now within a few benchmarks of parity on cost-optimized Chinese-lab hardware on vendor-selected evals — multimodal-agentic is where the enterprise-workflow revenue is; if the parity extends to independent eval, the migration axis becomes distribution and integration, not raw capability.
  1. V4-Flash-Vision-Exp Multimodal Variant Ships on DeepSeek API — Wins 3 of 11 Vendor-Selected Multimodal-Agentic Benchmarks vs Claude Opus 4.8 (August 22, 2026): DeepSeek on 2026-08-21 launched V4-Flash-Vision-Exp as an experimental multimodal variant of V4-Flash live on the DeepSeek API. On DeepSeek’s own leaderboard the model wins 3 of 11 agentic-multimodal benchmarks vs Claude Opus 4.8 and trails ~12 points on the hardest; Anthropic has not benchmarked back. Load-bearing framing to carry: do not lift Bloomberg’s “rivals” verb without the qualifier — the correct compact framing is “close to Opus 4.8 on 3 of 11 DeepSeek-selected multimodal benchmarks”; wait for third-party evaluation (Aider, LMSYS, LiveBench) before treating as a Chinese-lab parity result. Structural read: the multimodal-agentic axis — last generation’s US-lab moat — is now within a few benchmarks of parity on cost-optimized Chinese-lab hardware on vendor-selected evals. Multimodal-agentic is where enterprise-workflow revenue is; if the parity extends to independent evaluation, the migration axis becomes distribution and integration, not raw capability. Extends the V4-Flash family arc from the 2026-08-17-AI-Digest peak / off-peak repricing landing leg with a first multimodal capability variant leg — the V4-Flash tier is now no longer just a small-reasoning + open-weights + low-price benchmark reference; it’s also the first Chinese-lab launch onto the multimodal-agentic capability axis at a benchmarks-close-to-Opus-4.8 level. 30 / 60 / 90-day watch: independent third-party multimodal-agentic evals of V4-Flash-Vision-Exp against Opus 4.8 on non-DeepSeek-selected benchmarks; whether Anthropic benchmarks back on any of the 11 DeepSeek-selected multimodal evals; whether V4-Flash-Vision-Exp graduates from “Exp” and gets a formal pricing-page listing on the DeepSeek API; whether MiniMax, Kimi, or Qwen ship a comparable multimodal variant inside the same quarter (single-lab move vs China-frontier convergence test).

See also: DeepSeek, DeepSeek v4, DeepSeek V4 Pro, Claude Opus 4.8, AMD, MOC - AI Infrastructure, MOC - Open Source Models.