MODEL
Qwen
Overview
Qwen is Alibaba’s family of large language models spanning multiple scales and capabilities. The Qwen 3.5 series has demonstrated exceptional efficiency, with small models outperforming significantly larger competitors. Recent developments include the closed-source Qwen 3.6-Plus pivot, establishing Qwen as a major player in the LLM competitive landscape.
Timeline
- 2026-03-12-AI-Digest - Qwen 3.5 series initial tracking and comparison analysis
- 2026-03-16-AI-Digest - Qwen model performance updates and benchmark releases
- 2026-03-23-AI-Digest - Continued momentum in model refinement and evaluation
- 2026-03-27-AI-Digest - Performance benchmarking against competitor baselines
- 2026-03-31-AI-Digest - Model release updates and ecosystem integration
- 2026-04-03-AI-Digest - Qwen 3.6-Plus closed-source pivot announced; dethrones Llama on r/LocalLLaMA
- 2026-04-05-AI-Digest — Qwen 3.6 benchmarked by community against Gemma 4 and Llama 4 at comparable parameter counts; competitive performance confirmed.
- 2026-04-07-AI-Digest — Qwen3 base models mentioned in community context; Qwen3.6-Plus continues agentic coding focus
- 2026-04-07-AI-Digest — Qwen3 models in multiple sizes (0.6B–30B) released as open-source, continuing Alibaba’s open-weight strategy.
- 2026-04-09-AI-Digest — Qwen 3.5 cemented as one of the top two Apache 2.0 open-weights options on r/LocalLLaMA following Meta’s Muse Spark closed-source pivot. Community pragmatic consensus: Qwen 3.5 still wins on coding and tool calling, especially in
thinkingmode where extended chain-of-thought can be toggled per query, while Gemma 4 31B wins on multimodal, long-context, and structured output. Both fit cleanly on a 24 GB card at 4-bit quantization. - 2026-04-14-AI-Digest —
qwen-ai/qwen3-coder(128K context, tool calling) surfaces as a top April Hugging Face momentum project (2,800+ stars), cementing Qwen 3 Coder as the community default for code-specialist open-weights workloads. The broader r/LocalLLaMA consensus is a multi-model router pattern combining Qwen 3 Coder, Gemma 4 31B, DeepSeek V3, and Llama Stack. - 2026-04-18-AI-Digest — Qwen 3.5 enters week 2 of the GLM-5.1 vs Qwen 3.5 r/LocalLLaMA coding dispute, the second-most-active thread of the week. The pro-Qwen camp emphasizes broader language coverage, faster inference on commodity hardware, and a more mature tokenizer. The pro-GLM camp points to GLM-5.1’s SWE-Bench Pro lead (58.4%) and tool-use reliability on agentic loops. With Opus 4.7 extending the frontier-to-open-weights gap again (87.6% / 64.3% on the same benchmarks), the subtext is “which open model is the least-compromised local alternative” rather than “which open model is matching frontier.” Community working consensus: GLM-5.1 for agentic coding workflows, Qwen 3.5 for everything else, and run both if you have the VRAM. Qwen 3.5 holds position as the general-purpose open-weights default even as it cedes the narrow coding-specialist crown.
- 2026-05-10-AI-Digest — Top r/LocalLLaMA thread reports 80+ tok/sec at 80%+ draft acceptance running Qwen 3.6 35B A3B at 128K context (
-c 131072) on an RTX 4070 Super 12 GB, using the new multi-token-prediction PR againstllama.cppand theQwen3.6-35B-A3B-MTP-UD-Q4_K_XL.ggufquant (500 upvotes, 103 comments). Three of today’s r/LocalLLaMA top threads (this one, dual Mi50 MTP, the Q4_1 quants thread) thread the same MTP-on-modest-VRAM story — the PR is moving from experimental to default for the on-device crowd, with Qwen 3.6 35B A3B the marquee model. - 2026-05-17-AI-Digest — MTP support merges into llama.cpp (PR #22673, 12:06 UTC May 16) for the Qwen3.6 family; community benchmarks report up to +111% generation-speed gains on Qwen3.6 27B on AMD Strix Halo hardware (user-reported, not independently benchmarked). Separately, Qwen3.6-35B-A3B (3B active, MoE) lands on Terminal-Bench 2.0 leaderboard at 24.6% via the
little-coderscaffold, above Gemini 2.5 Pro on Gemini CLI (19.6%) but below Gemini 2.5 Pro on Terminus 2 (32.6%). - 2026-06-15-AI-Digest — Qwen surfaces on today’s HN community surface as the base-weights component in an alleged Rio de Janeiro “homegrown” LLM merge: a GitHub issue alleges Nex-N2 (marketed as a from-scratch Brazilian LLM) is in fact a merge — the Rio-3.5-Open-397B build appears to be ~0.6 Nex + 0.4 Qwen 3.5-397B-A17B, with weight fingerprints and tokenizer evidence on the thread. Another data point in the broader pattern of “sovereign-AI” launches being repackaged open weights — relevant to procurement attribution, vendor trust, and the open-weights-as-public-infrastructure thread. Pair with today’s Qwen 3.6 27B local-inference cross-reference for the local-first thread.
- 2026-06-24-AI-Digest — Qwen surfaces today via the “Qwen-AgentWorld: Language World Models for General Agents” paper (arXiv:2606.24597, ▲34) — language-based world models that simulate agentic environments across seven domains via extended reasoning chains, trained in three stages (capability injection, reasoning activation, reward-based refinement) on 10M+ real interaction trajectories, beating frontier baselines on the new AgentWorldBench. The corpus framing the digest carries: a usable simulator-plus-warm-up for agent RL from the Qwen team with open evaluation, hinting at a practical recipe for scaling general-purpose agents. Single paper on a community surface; the watch item is independent replication on AgentWorldBench by labs not affiliated with Alibaba.
Model Lineup
Qwen 3.5 Series
Efficient small models in the 0.8B-9B parameter range:
- 0.8B parameter model
- Mid-range variants
- 9B parameter model (top performer)
Qwen 3.6-Plus
- Architecture - Closed-source pivot from open-source strategy
- Parameter scale - Not publicly disclosed
- Release date - April 3, 2026
Key Specs & Benchmarks
Qwen 3.5 9B
- MMLU-Pro - 82.5 score
- GPQA Diamond - 81.7 score
- Significance - Beats models 13x larger in parameter count
Qwen 3.6-Plus
- SWE-bench - 78.8 score
- Competitiveness - Positioned against largest closed-source models
Market Position
Qwen 3.5’s performance on local LLM community forums (r/LocalLLaMA) has been so strong that it dethroned Llama as the community favorite, marking a significant shift in open-source model preferences. The transition to Qwen 3.6-Plus suggests Alibaba’s strategic pivot toward closed-source, service-based deployment models.
Timeline
- 2026-07-16-AI-Digest — Qwen serves as the on-device / LLM backbone for Apple Intelligence in mainland China after CAC added Apple’s generative AI stack to its approved-provider list, with Baidu supplying complementary capabilities. Commercial terms undisclosed; fall launch aligns with Apple’s OS cycle. Second time in ~45 days a Western frontier-model vendor has routed through a domestic Chinese model to reach the mainland market — Qwen is the anchor in Apple’s specific instance and the practitioner question flips from “can we launch in China” to “which domestic partner do we route through.” Same digest: PrismML‘s Bonsai 27B ships as a ternary / 1-bit quantised derivative of Qwen3.6-27B running on-device on an iPhone — Qwen is the substrate the on-device compression thread lands on rather than a first-party PrismML pretrain.
- 2026-07-17-AI-Digest — Qwen surfaces on two threads today. (1) Qwen’s Apple Intelligence China role sharpens into an explicit capability-routed architecture — Qwen handles language, Baidu handles visual (rather than the primary/secondary tier framing that had circulated earlier in the week); Alibaba ADRs closed +4.78% on the news, and Apple’s Q2 FY26 Greater China at $20.5B / +28% YoY makes the approval a material iPhone-upgrade lever going into September. Qwen is now the language anchor in a two-provider capability split that is the specific shape Beijing extracted as the price of CAC approval. (2) Qwen also appears in Bloomberg’s WAIC keynote setup as one of the Chinese lab families Xi’s remarks implicitly reference — Bloomberg names DeepSeek, Qwen, and Ant Group as having narrowed the frontier gap and won global open-weights adoption; the openness that makes them vectors for foreign intelligence use is what MIIT and CAC are actively consulting Alibaba, ByteDance, and Zhipu on restricting overseas access to top and unreleased open-weight models. Qwen is visibly positioned on both the inbound Apple/China distribution axis and the outbound MIIT/CAC restriction consultation axis inside the same news cycle.
- 2026-07-20-AI-Digest — The Qwen team previews Qwen 3.8 on Jul 19 — a 2.4T-parameter multimodal model launched as Qwen3-8-Max on the Qwen Cloud Token Plan at ~10% of standard-tier pricing. Marketing framing is “second only to Claude Fable 5” — Alibaba’s own positioning with no third-party benchmarks yet; MoE active-parameter count undisclosed. Weights are announced as forthcoming (“coming soon”); license is not disclosed. Right now the model is proprietary Max-Preview access on Alibaba’s cloud, with a top HN thread (Qwen 3.8, 819 pts / 569 cmts) whose linked tweet contains only a pricing-page link — the announcement is the artifact, not the shipment. Narrow read: unverified marketing until independent benchmarks land, and prior “Alibaba open-weight model coming soon” lines since Qwen 3.5 have shipped weights within 7–14 days. Structural read the digest carries: the Chinese open-weights cohort is now responding to Moonshot AI‘s Kimi K3 with a 72-hour counter-announcement — a distribution-competition cycle rather than a scheduled release cadence. OpenRouter Chinese-origin routed-token share extended from ~46% (2026-07-17-AI-Digest) to ~61% on the most recent third-party snapshot per the digest — distribution-majority thread still extending.