MODEL

Shieldstral

modeltopic-note

Overview

Shieldstral is a 3B-parameter Apache-2.0 safety-classifier model from Mistral, released Aug 4, 2026. Built on Ministral-3B plus a Pixtral image encoder; trained on 54.1M pairs across 12 languages; runs on a single 16GB GPU. Runtime-configurable yes/no prompts replace fixed content-policy taxonomies, making it policy-adaptive at inference time. Reported to match safety models roughly 7× its size on published benchmarks (gpt-oss-safeguard-scale). arXiv preprint at arXiv:2607.25857.

Timeline

  • 2026-08-06-AI-Digest — Coverage of Aug 4 release. Frames Shieldstral inside the thickening open-weight safety-tooling ecosystem alongside gpt-oss-safeguard; explicitly pushes back on the “open-weight capability closes but the safety gap widens” mainstream framing given yesterday’s UK AISI incident was attributed to closed frontier models.

Key Developments

  1. Release (Aug 4, 2026): Apache-2.0, 3B, 12 languages, 16GB GPU floor, runtime-configurable policy prompts, arXiv preprint alongside weights.