Mistral AI has released Shieldstral, a 3-billion-parameter open-weights safety classifier the company says matches or outperforms guardrail models up to seven times its size, under an Apache 2.0 licence.

Unlike traditional guardrail models that rely on a fixed taxonomy of harm categories baked into their weights, Mistral said Shieldstral accepts plain-language safety policies at inference time, letting a single checkpoint adapt to new moderation contexts without retraining. The model evaluates text and images through a unified question-answering format, reading a policy, a yes/no query, and the content itself, then returns a calibrated safety score rather than a discrete label.

Mistral said it evaluated Shieldstral against guardrail models up to seven times its size across four benchmark categories, text safety, refusal detection, policy adaptability, and multimodal safety, using evaluation samples held out from training. The company said the model runs on a single 16GB NVIDIA GPU and was trained on a mix of real and synthetic data spanning multiple existing safety taxonomies, unified into one training format. Mistral said it used contrastive training pairs, generated by rewriting safe text to violate specific policies, to teach the model to distinguish between similar but distinct policy boundaries, rather than memorising fixed categories.

Shieldstral was built on Forge, Mistral's platform for training and evaluating custom models. Mistral said it is an inaugural member of the Open Secure AI Alliance, alongside NVIDIA and other organisations, and plans further work on multilingual coverage and broader multimodal safety.


Who Owns AI Security in the Enterprise? Governance Is Still in Its Infancy
Who actually owns AI security in your organisation — and how mature is your governance around it? Two senior CISOs from vastly different environments give a straight answer: ownership sits with the CISO for now, and governance, even in well-run programmes, is still in its infancy. AI is shifting enterprise risk from defending infrastructure to defending decisions. Agentic AI operates semi- or fully autonomously, traditional security controls don’t fit probabilistic systems, and no single vendor covers the full attack surface. Speakers: Andy Holliday, CISO at Petrofac, Lester Godsey, CISO at Arizona State University and Stewart Tinson, Project Director, AI-360 You’ll learn: • Why the CISO is the only realistic owner of AI security risk for the next 5 years • Why agentic AI breaks deterministic security controls and what to do about it • How ASU built an actionable AI framework supporting 60+ large language models • Practical controls: API key hygiene, command whitelists, blast radius reduction • Why no single vendor can cover AI security end-to-end Key topics: Agentic AI risk • AI governance maturity • Threat model transformation • CISO ownership • Incident response for AI • Ethics & training data bias • Vendor landscape reality • Probabilistic vs deterministic controls For CISOs, CIOs, and risk leaders making decisions about AI adoption now.
Share this post
The link has been copied!