MODEL

Nemotron 3.5 Lightning

modeltopic-notenvidianemotronmoe

Overview

Nemotron 3.5 Lightning is a new variant in NVIDIA‘s open-weights Nemotron line surfaced on blogs.nvidia.com on Aug 11 2026 — a 30B MoE with ~3B active parameters, paired with a NeMo Switchyard routing / serving layer spanning RTX and DGX. Distinct from the earlier Nemotron 3 Series (Super / Ultra / Nano / VoiceChat / Omni), the Lightning variant is pitched at the low-active-parameter throughput envelope that matches recent hybrid open-weights entrants (Gated DeltaNet-2, Qwen3.6-35B-A3B) rather than the earlier dense-Nemotron cadence.

Timeline

  • 2026-08-12-AI-DigestNVIDIA pairs Nemotron 3.5 Lightning with the NeMo Switchyard routing / serving layer in an HN item on blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/. Model spec: 30B MoE / ~3B active. Switchyard positions the model inside NVIDIA’s own inference fabric spanning RTX (workstation / prosumer) and DGX (data-centre) tiers. Narrow read the corpus carries: model-plus-serving-fabric bundle, not a standalone weights drop — NVIDIA continues the pattern established by the 2026-07-08-AI-Digest Nemotron-Labs-Diffusion paper (SGLang / GB200 throughput as the reference figure) of shipping the model alongside its inference substrate rather than as bare weights. Structural read the corpus carries: NVIDIA continues bundling its own open-weights line with an inference fabric, pushing directly on the model-serving stack that AWS / Azure and independent hosts sell — the “who serves the model” competitive axis is now the visible NVIDIA-side move, not just accelerator supply. 30 / 60 / 90-day watch: whether independent benchmark posts land for the 30B/~3B-active configuration; whether Switchyard shows up in HF-hosted deployment references; whether the RTX-through-DGX serving-fabric framing gets picked up by any hyperscaler.

Key Developments

  1. 30B MoE / ~3B Active Variant Ships Paired With NeMo Switchyard Routing / Serving Layer (August 11, 2026): The Lightning variant places Nemotron inside the sparse-activation open-weights band (Gated DeltaNet-2, Qwen3.6-35B-A3B, sub-frontier MoE releases with <10B active) rather than the earlier dense-Nemotron 3 Series envelope. NeMo Switchyard is the serving-fabric partner, extending NVIDIA’s bundling of its open-weights line with its inference stack — the model release is inseparable from the serving-fabric announcement, and the RTX-through-DGX span is the demand-anchor framing.

  2. NVIDIA-as-Serving-Fabric-Vendor Continues to Extend Beyond Accelerator Supply: The Nemotron 3.5 Lightning + Switchyard bundle is the latest in a Q3 pattern (extending the 2026-07-08-AI-Digest Nemotron-Labs-Diffusion / SGLang / GB200 reference and the 2026-06-08-AI-Digest Naver–NVIDIA DSX + HyperCLOVA X Coalition roadmap) of NVIDIA shipping models plus their serving substrate as a single product rather than as separate hardware and model announcements — the competitive axis is now “who serves the model,” not only “who trains it.”

See also: Nemotron, NVIDIA, MOC - AI Infrastructure, MOC - Open Source Models.