MODEL
Qwen
Overview
Qwen is Alibaba’s family of large language models spanning multiple scales and capabilities. The Qwen 3.5 series has demonstrated exceptional efficiency, with small models outperforming significantly larger competitors. Recent developments include the closed-source Qwen 3.6-Plus pivot, establishing Qwen as a major player in the LLM competitive landscape.
Timeline
- 2026-03-12-AI-Digest - Qwen 3.5 series initial tracking and comparison analysis
- 2026-03-16-AI-Digest - Qwen model performance updates and benchmark releases
- 2026-03-23-AI-Digest - Continued momentum in model refinement and evaluation
- 2026-03-27-AI-Digest - Performance benchmarking against competitor baselines
- 2026-03-31-AI-Digest - Model release updates and ecosystem integration
- 2026-04-03-AI-Digest - Qwen 3.6-Plus closed-source pivot announced; dethrones Llama on r/LocalLLaMA
- 2026-04-05-AI-Digest — Qwen 3.6 benchmarked by community against Gemma 4 and Llama 4 at comparable parameter counts; competitive performance confirmed.
- 2026-04-07-AI-Digest — Qwen3 base models mentioned in community context; Qwen3.6-Plus continues agentic coding focus
- 2026-04-07-AI-Digest — Qwen3 models in multiple sizes (0.6B–30B) released as open-source, continuing Alibaba’s open-weight strategy.
- 2026-04-09-AI-Digest — Qwen 3.5 cemented as one of the top two Apache 2.0 open-weights options on r/LocalLLaMA following Meta’s Muse Spark closed-source pivot. Community pragmatic consensus: Qwen 3.5 still wins on coding and tool calling, especially in
thinkingmode where extended chain-of-thought can be toggled per query, while Gemma 4 31B wins on multimodal, long-context, and structured output. Both fit cleanly on a 24 GB card at 4-bit quantization. - 2026-04-14-AI-Digest —
qwen-ai/qwen3-coder(128K context, tool calling) surfaces as a top April Hugging Face momentum project (2,800+ stars), cementing Qwen 3 Coder as the community default for code-specialist open-weights workloads. The broader r/LocalLLaMA consensus is a multi-model router pattern combining Qwen 3 Coder, Gemma 4 31B, DeepSeek V3, and Llama Stack. - 2026-04-18-AI-Digest — Qwen 3.5 enters week 2 of the GLM-5.1 vs Qwen 3.5 r/LocalLLaMA coding dispute, the second-most-active thread of the week. The pro-Qwen camp emphasizes broader language coverage, faster inference on commodity hardware, and a more mature tokenizer. The pro-GLM camp points to GLM-5.1’s SWE-Bench Pro lead (58.4%) and tool-use reliability on agentic loops. With Opus 4.7 extending the frontier-to-open-weights gap again (87.6% / 64.3% on the same benchmarks), the subtext is “which open model is the least-compromised local alternative” rather than “which open model is matching frontier.” Community working consensus: GLM-5.1 for agentic coding workflows, Qwen 3.5 for everything else, and run both if you have the VRAM. Qwen 3.5 holds position as the general-purpose open-weights default even as it cedes the narrow coding-specialist crown.
- 2026-05-10-AI-Digest — Top r/LocalLLaMA thread reports 80+ tok/sec at 80%+ draft acceptance running Qwen 3.6 35B A3B at 128K context (
-c 131072) on an RTX 4070 Super 12 GB, using the new multi-token-prediction PR againstllama.cppand theQwen3.6-35B-A3B-MTP-UD-Q4_K_XL.ggufquant (500 upvotes, 103 comments). Three of today’s r/LocalLLaMA top threads (this one, dual Mi50 MTP, the Q4_1 quants thread) thread the same MTP-on-modest-VRAM story — the PR is moving from experimental to default for the on-device crowd, with Qwen 3.6 35B A3B the marquee model. - 2026-05-17-AI-Digest — MTP support merges into llama.cpp (PR #22673, 12:06 UTC May 16) for the Qwen3.6 family; community benchmarks report up to +111% generation-speed gains on Qwen3.6 27B on AMD Strix Halo hardware (user-reported, not independently benchmarked). Separately, Qwen3.6-35B-A3B (3B active, MoE) lands on Terminal-Bench 2.0 leaderboard at 24.6% via the
little-coderscaffold, above Gemini 2.5 Pro on Gemini CLI (19.6%) but below Gemini 2.5 Pro on Terminus 2 (32.6%). - 2026-06-15-AI-Digest — Qwen surfaces on today’s HN community surface as the base-weights component in an alleged Rio de Janeiro “homegrown” LLM merge: a GitHub issue alleges Nex-N2 (marketed as a from-scratch Brazilian LLM) is in fact a merge — the Rio-3.5-Open-397B build appears to be ~0.6 Nex + 0.4 Qwen 3.5-397B-A17B, with weight fingerprints and tokenizer evidence on the thread. Another data point in the broader pattern of “sovereign-AI” launches being repackaged open weights — relevant to procurement attribution, vendor trust, and the open-weights-as-public-infrastructure thread. Pair with today’s Qwen 3.6 27B local-inference cross-reference for the local-first thread.
- 2026-06-24-AI-Digest — Qwen surfaces today via the “Qwen-AgentWorld: Language World Models for General Agents” paper (arXiv:2606.24597, ▲34) — language-based world models that simulate agentic environments across seven domains via extended reasoning chains, trained in three stages (capability injection, reasoning activation, reward-based refinement) on 10M+ real interaction trajectories, beating frontier baselines on the new AgentWorldBench. The corpus framing the digest carries: a usable simulator-plus-warm-up for agent RL from the Qwen team with open evaluation, hinting at a practical recipe for scaling general-purpose agents. Single paper on a community surface; the watch item is independent replication on AgentWorldBench by labs not affiliated with Alibaba.
Model Lineup
Qwen 3.5 Series
Efficient small models in the 0.8B-9B parameter range:
- 0.8B parameter model
- Mid-range variants
- 9B parameter model (top performer)
Qwen 3.6-Plus
- Architecture - Closed-source pivot from open-source strategy
- Parameter scale - Not publicly disclosed
- Release date - April 3, 2026
Key Specs & Benchmarks
Qwen 3.5 9B
- MMLU-Pro - 82.5 score
- GPQA Diamond - 81.7 score
- Significance - Beats models 13x larger in parameter count
Qwen 3.6-Plus
- SWE-bench - 78.8 score
- Competitiveness - Positioned against largest closed-source models
Market Position
Qwen 3.5’s performance on local LLM community forums (r/LocalLLaMA) has been so strong that it dethroned Llama as the community favorite, marking a significant shift in open-source model preferences. The transition to Qwen 3.6-Plus suggests Alibaba’s strategic pivot toward closed-source, service-based deployment models.
Timeline
-
2026-07-16-AI-Digest — Qwen serves as the on-device / LLM backbone for Apple Intelligence in mainland China after CAC added Apple’s generative AI stack to its approved-provider list, with Baidu supplying complementary capabilities. Commercial terms undisclosed; fall launch aligns with Apple’s OS cycle. Second time in ~45 days a Western frontier-model vendor has routed through a domestic Chinese model to reach the mainland market — Qwen is the anchor in Apple’s specific instance and the practitioner question flips from “can we launch in China” to “which domestic partner do we route through.” Same digest: PrismML‘s Bonsai 27B ships as a ternary / 1-bit quantised derivative of Qwen3.6-27B running on-device on an iPhone — Qwen is the substrate the on-device compression thread lands on rather than a first-party PrismML pretrain.
-
2026-07-17-AI-Digest — Qwen surfaces on two threads today. (1) Qwen’s Apple Intelligence China role sharpens into an explicit capability-routed architecture — Qwen handles language, Baidu handles visual (rather than the primary/secondary tier framing that had circulated earlier in the week); Alibaba ADRs closed +4.78% on the news, and Apple’s Q2 FY26 Greater China at $20.5B / +28% YoY makes the approval a material iPhone-upgrade lever going into September. Qwen is now the language anchor in a two-provider capability split that is the specific shape Beijing extracted as the price of CAC approval. (2) Qwen also appears in Bloomberg’s WAIC keynote setup as one of the Chinese lab families Xi’s remarks implicitly reference — Bloomberg names DeepSeek, Qwen, and Ant Group as having narrowed the frontier gap and won global open-weights adoption; the openness that makes them vectors for foreign intelligence use is what MIIT and CAC are actively consulting Alibaba, ByteDance, and Zhipu on restricting overseas access to top and unreleased open-weight models. Qwen is visibly positioned on both the inbound Apple/China distribution axis and the outbound MIIT/CAC restriction consultation axis inside the same news cycle.
-
2026-07-20-AI-Digest — The Qwen team previews Qwen 3.8 on Jul 19 — a 2.4T-parameter multimodal model launched as Qwen3-8-Max on the Qwen Cloud Token Plan at ~10% of standard-tier pricing. Marketing framing is “second only to Claude Fable 5” — Alibaba’s own positioning with no third-party benchmarks yet; MoE active-parameter count undisclosed. Weights are announced as forthcoming (“coming soon”); license is not disclosed. Right now the model is proprietary Max-Preview access on Alibaba’s cloud, with a top HN thread (Qwen 3.8, 819 pts / 569 cmts) whose linked tweet contains only a pricing-page link — the announcement is the artifact, not the shipment. Narrow read: unverified marketing until independent benchmarks land, and prior “Alibaba open-weight model coming soon” lines since Qwen 3.5 have shipped weights within 7–14 days. Structural read the digest carries: the Chinese open-weights cohort is now responding to Moonshot AI‘s Kimi K3 with a 72-hour counter-announcement — a distribution-competition cycle rather than a scheduled release cadence. OpenRouter Chinese-origin routed-token share extended from ~46% (2026-07-17-AI-Digest) to ~61% on the most recent third-party snapshot per the digest — distribution-majority thread still extending.
-
2026-08-01-AI-Digest — Alibaba‘s Qwen team publishes Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents (arXiv:2607.28227, ▲271 on Hugging Face) laying out a foundation-model roadmap for computer-use agents targeting reliable operation on real devices, cross-platform workflows, hybrid GUI+CLI execution, long-horizon tasks, and autonomous self-improvement. Narrow read: research paper on the community-surface pass; roadmap frames future model direction rather than a released Qwen-UI-Agent product. Structural read the corpus carries: positions Qwen to compete directly with Anthropic and OpenAI‘s computer-use agent stacks on an open-weights posture — the GUI-agent lane where Anthropic’s Claude for Chrome and OpenAI’s Operator have been the main frontier-lab entrants now has an open-weights research anchor with a stated foundation-model roadmap. 60-day watch: whether the paper’s roadmap converts to a released Qwen-UI-Agent checkpoint on Hugging Face, and whether independent OSWorld-style benchmarks land it above, at, or below Claude Opus 4.8 and GPT-5.6 Sol on browser-and-desktop agent tasks.
-
2026-08-02-AI-Digest — Qwen surfaces today via the Qwen-UI-Agent Technical Report landing on Hugging Face at ▲278 (arXiv:2607.28227) — Alibaba’s foundation GUI agent unifies mobile / computer-use / web / DeepSearch, interleaves GUI operations with CLI execution, and trains via online RL on 100+ turn trajectories across 10,000 concurrent environments. Reported benchmarks: 82.1% MobileWorld, 79.5% OSWorld-Verified, 73.6% WebArena — matching or beating Claude Opus 4.8 / Gemini 3.1 Pro / GPT-5.6 Sol on the reported benches. Narrow read: strongest open-weights GUI agent to date on reported numbers, though these are Alibaba’s own bench results on their own paper — independent OSWorld-style replication is what would move this from co-emergence to a genuine open-weights GUI-agent frontier. Structural read the corpus carries: concrete template for how frontier labs are industrialising agent training at commodity-environment scale — 10,000 concurrent environments with 100+ turn trajectories is training-recipe scale that would previously have been a frontier-lab-only capability, and the corpus should log this as the Qwen (Alibaba) foundation-model-of-the-day on the community-surface pass. Extends yesterday’s roadmap-paper entry with the benchmark-numbers reported beat, still on the same underlying arXiv report.
-
2026-08-15-AI-Digest — Alibaba drops Qwen 3.8 27B as a mid-size FP8 checkpoint straight to Hugging Face under Apache 2.0 — HN thread at 995 pts / 642 cmts. Sits alongside the Aug 12 Qwen3.8-2.4T-A95B frontier ship as the mid-size sibling on the same Qwen 3.8 release line. Narrow read the digest carries: mid-size Qwen releases keep resetting the local/inference-cost bar and get adopted into the OSS stack within hours; no vendor pitch, HN commentary is doing the initial evaluation work. Structural read: the Qwen 3.8 line now spans frontier-MoE (2.4T-A95B) and mid-size dense/FP8 (27B) inside a three-day window, extending the frontier-MoE-AND-small-dense bifurcation the 2026-08-13-AI-Digest entry flagged with Alibaba anchoring both poles on its own release line rather than ceding the small-dense pole to Meta / distillation labs. Same digest’s Aider
[!note]: Qwen 3.8 releases do not yet have Aider entries; the polyglot-freeze reading against fresh Qwen 3.8 momentum is a leaderboard-lag artifact, not a capability signal. -
2026-09-01-AI-Digest — “On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability” — a research paper from the Qwen team on the next-generation architecture (arXiv:2608.30320, ▲3). Qwen3.8-Flash-Next is a 125B-param sparse MoE with 6B activated and 51B off-accelerator n-gram embeddings, using Gated DeltaNet plus selective full attention to match a larger predecessor at ~1/9 the training FLOPs. Load-bearing framing the digest carries: concrete architecture-paper for the next Qwen generation, with hybrid attention and host-memory embeddings that push efficiency without a capability cliff. Narrow read: research paper on the community-surface pass; single-vendor claim on training-FLOP efficiency until independent evaluations land. Structural read: extends the Qwen team’s public architectural-roadmap posture (following the Qwen-UI-Agent Technical Report landing in early August) with a training-efficiency axis on the next-gen Flash line. Log against MOC - Open Source Models and MOC - AI Infrastructure.
-
2026-09-07-AI-Digest — Qwen surfaces today as the base backbone in a new omni paper — “Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue” (arXiv:2609.04250, ▲12). A Qwen-2.5-7B-based omni model emits speech together with facial, hand, and body motion from shared hidden states, replacing the speech-then-motion cascade. Matches teacher-cascade motion quality within ~2% while running 5.4x faster (RTF 0.78) at 2.62% WER. Narrow read: another external paper picking Qwen-2.5-7B as the small-open-weights substrate for a novel real-time embodied-avatar stack. Structural read: Qwen 2.5-7B remains the go-to open backbone for adjacent research pipelines even as the Qwen 3.8 line dominates the coverage cycle — the base’s install-base advantage compounds when downstream researchers pick it as the recipe substrate. Points at a viable real-time embodied-avatar path without paying for two inference passes. Log against MOC - Open Source Models.
-
2026-08-30-AI-Digest — Alibaba ships Qwen 3.8-Flash — a lower-cost tier positioned against Anthropic‘s Claude Opus 5 on the flagship axis (~30× discount) and DeepSeek V4-Flash on the low-cost axis (Bloomberg). Public pricing: $0.16 / M input · $0.47 / M output vs DeepSeek V4-Flash’s $0.14 / $0.28. Narrow read the digest carries: Bloomberg’s directional framing against DeepSeek is misleading — Qwen 3.8-Flash is priced at the V4-Flash tier, slightly above on both dimensions, not below; it joins that tier, it does not undercut it. Structural read: do NOT extend to “Chinese labs relentlessly compress token prices further” — the compression from Claude-tier to Flash-tier already happened in Q2; Qwen 3.8-Flash is Alibaba entering the existing floor. Correct frame: the Flash-tier pricing band is now crowded with three credible open-weight-adjacent options (DeepSeek V4-Flash, Qwen 3.8-Flash, Hy4 Preview‘s $0.83/$2.50 flagship-lite tier) — differentiator is capability profile and licence, not price; read the capability claim against Aider polyglot’s still-all-US top-3, not against Bloomberg’s “outperforms” shorthand. Extends the Qwen 3.8 release line (Aug 3 Max preview → Aug 12 2.4T-A95B → Aug 15 27B FP8 → Aug 30 Flash low-cost tier) with the fourth SKU on the same three-week roll. Log against MOC - Open Source Models.