MODEL
Leanstral 1.5
Overview
Leanstral 1.5 is Mistral‘s July 2026 formal-math / theorem-proving model, announced on 2026-07-04 and framed around theorem-proving accessibility (“Proof abundance for all”). Positions Mistral against DeepSeek-Prover and Kimina in the formal-math / proof-assistant lane, locating Mistral’s incremental releases outside the mainstream-benchmark race.
Timeline
- 2026-07-04-AI-Digest — Mistral announces Leanstral 1.5 (“Proof abundance for all”) on the Mistral news page; HN front page at 138 pts / 35 cmts. Framed around theorem-proving accessibility rather than mainstream reasoning benchmarks — pushes further into the formal-math / proof-assistant lane where DeepSeek-Prover and Kimina have been setting the pace, and locates Mistral’s incremental releases outside the mainstream-benchmark race.
- 2026-07-06-AI-Digest — Benchmark and bug-catching numbers land: Apache-2.0, 119B-total / 6B-active MoE, 100% on miniF2F, 587 of 672 on PutnamBench, tops FATE-H (87) and FATE-X (34) on the open-source field, and — during evaluation — surfaced five previously unknown bugs across 57 open-source repositories (including a
varintegeroverflow in a Rust codebase). Narrow read: open-source SOTA on Lean 4 formal-math benchmarks with demonstrable transfer to code verification on real projects. Structural read the digest carries: extends the “open-weights closing the gap on closed baselines” thread the corpus has been tracking through Reflection and Apertus releases, but on a formal-verification benchmark where DeepMind’s AlphaProof-class systems remain off-benchmark and non-comparable. The 60-day test the digest holds: whether the “5 real bugs” number is reproduced by an independent adopter — that separates novel evaluation datum from shipping-product-category signal.
Key Developments
-
Formal-Math Positioning (July 4, 2026): Leanstral 1.5 lands as a theorem-proving-accessibility release rather than a general-reasoning model — a discipline-specific Mistral entry into the formal-math lane where DeepSeek-Prover and Kimina have been the pacesetters. The disciplined framing worth carrying: Mistral is using formal-math as a differentiation lane outside the closed-frontier benchmark race, not as a claim on frontier reasoning capability.
-
Open-Source SOTA on Lean 4 + Five Real OSS Bugs (July 6, 2026): 119B-total / 6B-active MoE under Apache-2.0 hits 100% on miniF2F, 587/672 on PutnamBench, tops FATE-H (87) and FATE-X (34) on the open-source field, and during evaluation surfaced five previously unknown bugs across 57 open-source repositories (including a
varintegeroverflow in a Rust codebase). The formal-verification benchmark cohort keeps DeepMind’s AlphaProof-class systems off-benchmark and non-comparable, so the “open-weights closing on closed baselines” thread applies specifically on the formal-math axis rather than on the general-reasoning axis. 60-day follow-on test: independent reproduction of the five-bugs number by a non-Mistral adopter (2026-07-06-AI-Digest).
Related
See also: Mistral, MOC - Open Source Models.