Introducing Shieldstral.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.
Follow Mistral AI Blog to make it a durable For You signal.
Mistral AI has released Shieldstral, a 3-billion-parameter open-weights multimodal safety classifier that frames content moderation as a policy-adaptive question-answering task. Instead of relying on a fixed set of harm categories, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. The model matches or outperforms open guard models up to seven times its size on text safety benchmarks and sets a new state of the art on multimodal moderation, while running efficiently on a single 16GB NVIDIA GPU. Released under the Apache 2.0 license, Shieldstral enables developers to specify and adjust safety guardrails on the fly, making it easier to tailor moderation to different products, audiences, and contexts without retraining.