MODEL
DeepSeek V4.1-Flash
modeltopic-note
Overview
DeepSeek V4.1-Flash is a 552B-parameter MoE model with ~8B active per token, MIT-licensed weights on HuggingFace, native multimodal (image-in, text-out), and a 1M-token context. Trained on 45T mixed text-and-image tokens; priced at $0.15/M input and $0.60/M output (off-peak) on the official API with cached-input as low as $0.003/M. A MoE-style refresh of the V4 family rather than a novel architecture — the practitioner takeaway is memory-and-cost efficiency for agent inference loops. See also the earlier DeepSeek-V4-Flash for the prior V4-family Flash variant.
Timeline
- 2026-09-11-AI-Digest — DeepSeek released V4.1-Flash on Sep 10 — 552B MoE / ~8B active, MIT weights on HuggingFace, 1M context, multimodal, 45T tokens. Off-peak pricing $0.15/M input, $0.60/M output; cached-input as low as $0.003/M. DeepSeek’s own benchmark table claims narrow leads on Terminal-Bench 2.1, DeepSWE, AutomationBench, ALE, and CyberGym against GPT-5.6 Sol, Claude Fable 5.1, and Opus 5 — but no LMArena or Simon Willison independent evaluation has yet surfaced, so the frontier-parity claim rides on DeepSeek-selected agentic-coding benchmarks and off-peak cached-input rates. Carry as frontier-grade agent-inference-cost floor lowers again on published DeepSeek evals — awaiting independent verification.
- 2026-09-14-AI-Digest — The API cutover landed at 04:00 UTC today — every DeepSeek V4-Pro call now reroutes to V4.1-Flash at Flash pricing until V4.1-Pro ships. Bloomberg Intelligence’s per-token math puts the effective cut at as much as 32% (off-peak output at $0.60/M tokens, cached input as low as $0.003/M, peak rates 2×), reversing the August V4-Pro hike that followed DeepSeek’s coding-agent-rival launch. MiniMax and Z.ai fell >8% in Hong Kong on the news; Alibaba closed -2%. This is the operational-cost side of the 2026-09-11-AI-Digest release — the model didn’t change today, but the price floor for anyone still paying V4-Pro rates actually collapses at 04:00 UTC. The Aider polyglot top-5 still doesn’t list a DeepSeek entry inside the 88.0–81.3% band gpt-5 and o3-pro occupy — the frontier-parity claim rides on DeepSeek-selected agentic-coding benchmarks and off-peak cached-input rates until independent evals surface. Reframe worth carrying:
frontier-grade agent-inference-cost floor keeps lowering on DeepSeek-selected benchmarks and off-peak cached-input rates, notV4.1-Flash is now the price-quality frontier. Log against MOC - Open Source Models and MOC - AI Infrastructure.
Key Developments
- V4.1-Flash release (Sep 2026): 552B MoE with ~8B active per token, MIT weights, 1M context, native multimodal — a MoE-style refresh of the V4 family, not a novel architecture. Trained on 45T tokens. Pricing lands memory-and-cost efficiency for agent inference loops as the load-bearing practitioner takeaway; independent third-party benchmark corroboration still pending on Claude Fable / Opus 5 comparisons.