MODEL

Ring-2.5-1T-Zero

modeltopic-notechinese-aiopen-sourcereinforcement-learning

Overview

Ring-2.5-1T-Zero is a 1-trillion-parameter language model from Ant Group and Renmin University, published on arXiv on 2026-07-15 (arXiv:2607.12395) as the largest publicly disclosed model trained with pure zero-supervision reinforcement learning (no supervised-fine-tuning stage). The paper reports emergent structured reasoning, self-verification, and parallel-reasoning behaviors on math benchmarks. Released the same week the Chinese-open-weight distribution-majority story surfaced at the Hugging Face and OpenRouter aggregator level, Ring-2.5-1T-Zero is the training-recipe signal one axis further left than the distribution-side news.

Timeline

  • 2026-07-15-AI-Digest — Paper posted to arXiv (arXiv:2607.12395) under the title “Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning.” Abstract reports emergent structured reasoning, self-verification, and parallel-reasoning behaviors on math benchmarks. Largest publicly disclosed pure-RL post-training result to date. Comes from a Chinese lab, in the same news cycle as the Chinese-open-weight distribution-majority story — temporal shape is coordinated whether or not intent is.

Key Developments

  1. Largest Publicly Disclosed Pure-RL Post-Training Run (July 15, 2026): 1T parameters trained end-to-end with zero-supervision RL (no SFT stage). Abstract-level claims of emergent structured reasoning, self-verification, and parallel-reasoning behaviors on math benchmarks — pending independent replication. Positions Ant Group and Chinese labs on the training-recipe axis alongside their compounding on the distribution axis (Hugging Face 41% of spring downloads, top-6 OpenRouter sweep).
  • 2026-07-16-AI-DigestRing-Zero paper “Scaling Zero RL to a Trillion Parameters for Emergent Reasoning” (arXiv:2607.12395, ▲45) — cadence continuation on today’s HN “Papers” section, one day after the Jul 15 first-day arXiv drop. The paper describes the Zero-RL training pipeline (no human-annotated data) scaled to 1T parameters with clipped importance sampling and training-inference ratio correction as the stabilization tricks; distinct discovery and sharpening phases surface where models spontaneously develop structured formatting, self-verification, and parallel reasoning. Corpus framing: first public demonstration that pure-RL reasoning training keeps paying off at trillion-parameter scale. Same-week temporal shape with the Chinese-open-weight-distribution-majority story (2026-07-15-AI-Digest) is coordinated in shape whether or not it is coordinated in intent.

See also: Ant Group, MOC - Open Source Models.