MODEL
Kimi K2.5
modeltopic-note
Overview
Kimi K2.5 is a 1-trillion-parameter mixture-of-experts model. In May 2026 it became notable as the subject of the first documented LLM inference build using Intel Optane Persistent Memory, demonstrating local inference at frontier scale on unusual prosumer hardware.
Timeline
- 2026-05-12-AI-Digest — An r/LocalLLaMA post (517 points) documented the first known LLM inference build using Intel Optane Persistent Memory (DIMM-form-factor non-volatile memory between DRAM and SSD on the latency curve) to run Kimi K2.5 locally at over four tokens per second. Optane has been EOL since 2022, so the parts pool is fixed and shrinking, but the build demonstrates a memory tier that materially expands the addressable working-set for very large models on commodity hardware — contributing to the MTP-and-memory-tier moment in the local inference community.
Key Developments
- Optane PMem Inference Build: First documented use of Intel Optane Persistent Memory for LLM inference — running a 1T-parameter MoE model at 4+ tok/s on prosumer hardware, showing that non-standard memory tiers can unlock frontier-scale local inference even with limited GPU memory.
- 2026-07-16-AI-Digest — KnowAct-GUIClaw paper (arXiv:2607.12625, ▲29) uses the Kimi-2.6 open-weights build as the base model for a Know-Route-Act-Reflect GUI-agent framework that hits 64.1% on MobileWorld, beating Seed-2.0-Pro and GPT-5.5. Reference-substrate signal rather than fresh Kimi first-party news — the Kimi lineage is now the open-weights anchor for a personal-GUI-assistant benchmark leader against closed models. Corpus framing: open-weights GUI-agent stack overtaking closed models on a long-horizon benchmark reinforces that memory + skill libraries — not raw base models — are becoming the differentiator.
- Kimi-2.6 as Base Substrate for KnowAct-GUIClaw MobileWorld Leader (July 16, 2026): The Know-Route-Act-Reflect GUI-agent framework built on Kimi-2.6 hits 64.1% on MobileWorld, beating Seed-2.0-Pro and GPT-5.5. Kimi lineage is the open-weights anchor for a closed-model-beating benchmark on a long-horizon GUI task — reinforces the frame that memory + skill libraries are becoming the differentiator over raw base models.
- 2026-08-09-AI-Digest — Kimi K2.5 surfaces today as the counter-example pricing anchor against DeepSeek‘s Aug 6 developer-email warning of a substantial cross-service API price hike — the digest reads “Kimi K2.5 sits at $0.60 / $3 on Moonshot AI‘s own pricing page” alongside “Qwen 3.5 Flash still lists at $0.10 / $0.40 per M” as the load-bearing evidence that DeepSeek’s move is DeepSeek-specific, not a China-wide pivot toward profitability. Also: current prices in the DeepSeek pricing-adjustment story explicitly compare DeepSeek V4 Flash at $0.14 / $0.28 versus K2.5 at $3 / $15 as the frontier-tier comparator — K2.5 anchors the upper end of the Chinese-open-weight commodity band the DeepSeek move is being priced against. Narrow read: K2.5 is not the subject of today’s news but the pricing-durability comparator on the upper end of the Chinese-open-weight band. Structural read: K2.5 continues to hold as the Moonshot AI flagship-tier pricing anchor alongside Kimi K3 — the “just use DeepSeek” default weakens specifically, not the broader China open-weight low-cost story. 30 / 60 / 90-day watch: whether Qwen or Kimi K2.5 follow DeepSeek with matching hikes (that would upgrade the framing to a real China pricing pivot).