MODEL

Kimi K3

modeltopic-noteopen-sourcechinese-ai

Overview

Kimi K3 is Moonshot AI‘s flagship open-weights mixture-of-experts model, released 2026-07-17 at roughly 2.8T total parameters with a 1M-token context window and pricing set at $3 per M input / $15 per M output (with a $0.30/M cache-hit discount) — the same headline pricing as Anthropic‘s Claude Sonnet 5 and materially below the $5/$25 of Claude Opus 4.7. Active-parameter count is not disclosed, which matters for cost-per-throughput reads against Inkling‘s 41B active. Positioned as a frontier-adjacent open entrant on the distribution axis rather than the capability-ceiling axis.

Timeline

  • 2026-07-17-AI-DigestMoonshot AI released Kimi K3 — 2.8T MoE with 1M context at $3/$15 per M tokens; dominant HN discussion at the top of the front page today. Simon Willison’s release-day post is careful about benchmark framing: pelican-style microbenchmarks are saturated at the frontier and the honest test for K3 is agentic tool-calling and long-conversation reliability, not one-shot SVG generation. Same digest: the Aider polyglot top-5 remains unchanged from yesterday (K3 not yet scored) — a snapshot benchmark that materially trails the open-weight release cycle. Narrow read: pricing is the story, not raw scale — a claimed 3T-class open model at GPT-5.4 tier undercuts Opus 4.7 output by ~40%. Structural read the corpus carries: the two-leaderboards frame from earlier this week now has a fresh price point on the distribution-share axis — OpenRouter telemetry shows Chinese-origin models at ~46% of routed tokens vs US ~30% (down from ~70% in June ‘25), and K3 at Sonnet pricing is the kind of drop that accelerates that mix. 60-day watch: K3’s Aider polyglot entry once submitted — a top-5 finish at Sonnet pricing would collapse the “cheap but weaker” default assumption; a lower placement re-anchors the price/performance-per-tier read.
  • 2026-07-18-AI-DigestKimi K3 named by Bloomberg as one accelerant of the chip-stocks bear-market entry — SOX widening its drop from the late-June record to ~20% — alongside Samsung soft prelims and the second Netlist ITC probe, but the digest holds the spark-on-dry-tinder framing: SOX had already shed ~7% on July 7 Samsung prelims and Applied Materials had shed ~10% before K3 shipped; TNW literally frames the rout as “already loaded” when K3 landed. Two benchmark corrections carry: K3 beats Claude Opus 4.8 and GPT-5.5 but trails Claude Fable 5 and GPT-5.6 Sol on coding benchmarks per VentureBeat — one notch below “Fable 5 tier” rather than the “substantially outperform” framing of early Bloomberg coverage. Weight availability asterisked: MXFP4-quantized weights arrive 2026-07-27, not launch, and full-precision self-hosting still requires ~1.4 TB of storage and 8–16 nodes of 8×H100/B200 (~$80K in DGX Spikes at full precision) — “downloadable and cheap” is API-cheap in practice, not median-practitioner-downloadable. Aider polyglot top-5 (fetched Jul 18) still shows K3 absent; a top-5 finish once submitted would collapse the “cheap but weaker” default.
  • 2026-07-19-AI-Digest — Kimi K3 surfaces today as the Sonnet-tier API-pricing anchor ($3/$15 per M) against which Anthropic‘s Claude Fable 5 subscription cuts are being measured: with Max/Team Premium capped at 50% of already-reduced weekly caps and Pro/Team Standard losing bundled access with a one-time $100 API credit and then paying list ($10/$50), the price-per-throughput comparison shifts materially in the open-weights direction at the Pro-tier practitioner segment specifically — Pro subscribers pushed to Fable 5 API rates now weigh K3-at-Sonnet-pricing on the same axis. Also contextually adjacent to today’s UK AISI open-weight cyber-capability gap compression (6–10mo → 4–7mo against frontier) — the 2026-07-15-AI-Digest distribution-share thread continues to compound on the capability axis as well as the price axis. K3’s coding-benchmark asterisk from 2026-07-18-AI-Digest holds unchanged: still beats Claude Opus 4.8 and GPT-5.5, still trails Claude Fable 5 and GPT-5.6 Sol; MXFP4 weights arrive 2026-07-27, not landed today.
  • 2026-07-20-AI-Digest — Kimi K3 surfaces today as the release that opened the same-week China-open-weights counter-announcement cycle — the digest frames Alibaba‘s Qwen 3.8 preview as “the second China-open-weights response to Kimi K3 in 72 hours,” positioning K3’s Jul 17 launch as the trigger for the 72-hour cadence Alibaba’s preview compresses against. K3 remains named in the day’s coverage as the retained Sonnet-tier open comparator against Claude Fable 5 on the Pro-tier cutover story. Aider polyglot top-5 is now on its fifth consecutive day with identical rows and percentages — K3 still absent from the board (inclusion-lag traces back to 2026-06-12-AI-Digest) alongside Claude Fable 5 and GPT-5.6 Sol. No fresh Moonshot/K3 product action today; the corpus logs today as trigger + comparator framing.
  • 2026-07-21-AI-DigestKimi K3 lands at the centre of today’s digest as the Bloomberg-headlined market anxiety — K3 shipped 2026-07-16 as a 2.8T-parameter open-weight at $3 / $15 per M tokens ($0.30 cached input), identical to Sonnet 5‘s post-Sept 1 rate card and ~6× the K2.6 rate of $0.95 / $4. The Bloomberg framing centres market anxiety and DeepSeek-style reevaluation of US-lab compute moats; the digest reframes with the disciplined read that the pricing move underneath is the actual story, not the parameter count. The “China ships cheap open weights” thread from earlier in the corpus (see 2026-04-15-AI-Digest, 2026-06-02-AI-Digest) has now inverted for at least this release — K3 is priced at Sonnet-parity, not below it. Structural read: the pattern is not “China open-weights are winning” as a single-winner story — it’s a split. Combined Chinese providers hold >45% of OpenRouter weekly-token share on the inference-volume battlefield, but Anthropic and OpenAI still hold enterprise-integration and regulated-workload battlefields intact. K3’s Sonnet-parity pricing is Moonshot moving off the inference-volume playbook into the enterprise-margin one, not the other way around.
  1. Sonnet-Parity Pricing as Enterprise-Margin Move Rather Than Inference-Volume Play (July 21, 2026): K3 at $3/$15 per M tokens ($0.30 cached input) is ~6× the K2.6 rate of $0.95/$4 and identical to Sonnet 5’s post-Sept 1 rate card. That inversion of the earlier “China ships cheap open weights” thread — same lab, same category, priced up not down — is the load-bearing signal. Combined Chinese providers >45% OpenRouter weekly-token share on inference-volume, but US closed labs keep the enterprise-integration and regulated-workload lanes. Read the K3 pricing move as Moonshot rebalancing from the inference-volume battlefield to the enterprise-margin battlefield rather than a single-winner “China wins” framing.

Key Developments

  1. Sonnet-Tier Open Frontier Pricing Entry (July 17, 2026): 2.8T MoE with 1M-token context, $3/$15 per M tokens plus $0.30/M cache-hit discount — same headline pricing as Claude Sonnet 5 and ~40% below Claude Opus 4.7 output. Active-parameter count undisclosed matters for cost-per-throughput reads. The load-bearing corpus signal: this is the first open model to price directly on top of the closed commodity tier at the ceiling of size claims, and it lands the same news cycle the aggregator-level Chinese-origin distribution-share majority becomes visible on OpenRouter (~46% vs US ~30%). Puts serious pressure on the commodity-tier bracket without touching the frontier reasoning ceiling.

See also: Moonshot AI, Kimi K2.5, Claude Sonnet 5, Claude Opus 4.7, Inkling, MOC - Open Source Models.