Source-linked AI summary

Shieldstral

Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli, Guillaume Lample, Maarten Buyl, Maximilian Augustin, Maximilian Müller, Pierre Stock, Tom Bewley, Wassim Bouaziz, Yimu Pan, Abdelaziz Bounhar, Abhijeet Somani, Aditi Kabra, Adrian Valente, Adrien Petralia, Adrien Sadé, Alan Jeffares, Albert Jiang, Aleksandr Timashov, Alexandre Cahill, Alexandre Gavaudan, Alexandre Laval, Alexandre Sablayrolles, Amélie Héliou, Amos You, André Jonasson, Andrew Bai, Andrew Ehrenberg, Andrew Zhao, Angele Lenglemetz, Anmol Agarwal, Arata Suzuki, Arjun Majumdar, Arthur Fournier, Artjom Joosen, Aylin Guliz Akkus, Aysenur Karaduman, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Benjamin Holzschuh, Benjamin Lefaudeux, Benjamin Tibi, Bernhard Stadlbauer, Błażej Osiński, Camille Le Scao, Chaoran Yu, Charlotte Cronjäger, Chen-Yo Sun, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Corentin Barreau, Corentin Sautier, Cristiana-Diana Diaconu, Cyprien Courtot, Daniel Marczak, Darius Dabert, Diego de Las Casas, Dominik Nuss, Dylan Rubini, Dzmitry Soupel, Elizaveta Demyanenko, Elliot Chane-Sane, Emilien Fugier, Emmanuel Gottlob, Erik Aas, Etienne Goffinet, Étienne Millon, Eujeong Choi, Fabian Paischer, Fabian Schlager, Faruk Ahmed, Federico Baldassarre, Filip Szatkowski, Florian Wiesner, Gabrielle Berrada, Gaëtan Ecrepont, Gaétan Lepage, Gaspard Blanchet, Gaspard Donada-Vidal, Gauthier Delerce, Gauthier Guinet, Genevieve Hayes, Georgii Novikov, Gianluca Galletti, Guillaume Breton, Guillaume Kunsch, Guillaume Martin, Guillaume Raille, Gunjan Dhanuka, Gunshi Gupta, Han Zhou, Harshil Shah, Hasan Furkan Vural, Hédi Hadiji, Hope McGovern, Hugo Cisneros, Hugo Thimonier, Indraneel Mukherjee, Ivan Cuevas Salazar, Jacques Sun, Jan Ludziejewski, Jason Rute, Jean Quentin, Jean-Hadrien Chabran, Jean-Malo Delignon, Jie Zhang, Joachim Studnia, Joep Barmentlo, Johannes Brandstetter, John Harvill, Jonas Amar, Jonas Schweizer, Joséphine Delas, Josselin Somerville, Julien Denize, Julien Tauran, Kartik Khandelwal, Khyathi Raghavi Chandu, Kilian Tep, Kush Jain, Larissa Laich, Laura Calem, Laurence Aitchison, Laurent Callot, Laurent Fainsin, Léo Cotteleer, Léonard Blier, Lingxiao Zhao, Louis Martin, Louis Serrano, Lucile Saulnier, Ludovic Ho Fuh, Luis Montero, Manon Chossegros, Marcin Możejko, Margaret Jennings, Markus Hennerbichler, Martin Alexandre, Mathieu Poirée, Mathieu Schmitt, Mathilde Guillaumin, Matthieu André, Matthieu Dinot, Matthieu Futeral, Maurits Bleeker, Mauro Comi, Max Mynter, Maxim Berman, Maxime Darrin, Maxime Louis, Melina Jingting Laimon, Mert Unsal, Mia Chiquier, Michael Pilcer, Michał Pietruszka, Michał Zając, Mikhail Biriuchinskii, Minh-Quang Pham, Minwoo Kang, Morgane Rivière, Namit Katariya, Nathan Grinsztajn, Nathan Simpson, Neeraj Aggarwal, Neha Gupta, Ola Mysiak, Oliver Leicht, Olivier Bousquet, Olivier Duchenne, Parag Jain, Patricia Wang, Patrick Blies, Patrick von Platen, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Pavan Kumar Reddy, Pavel Kuksa, Philippe Pinel, Philomène Chagniot, Pierre-André Savalle, Piotr Milos, Prateek Gupta, Pravesh Agrawal, Quentin Desreumaux, Quentin Torroba, Quercus Hernandez, Ram Ramrakhya, Randall Isenhour, Ranjit Parva, Raul Perez Pelaez, Reinhard Sonnleitner, Rémi Delacourt, Richard Kurle, Rishi Shah, Rob Romijnders, Rohin Arora, Romain Sauvestre, Roman Soletskyi, Rosalie Millner, Rupert Menneer, Sagar Vaze, Samuel Barry, Samuel Belkadi, Samuel Humeau, Sanchit Gandhi, Sandeep Subramanian, Sarthak Mittal, Saskia Adaime, Sean Cha, Sebastian Kaltenbach, Shashwat Dalal, Shashwat Verma, Sherif Waly, Shrimai Prabhumoye, Siddhant Waghjale, Siddharth Gandhi, Simon Lepage, Simon Sorg, Soham Ghosh, Sophie Marbach, Srijan Mishra, Stanislas Lange, Steve Hong, Sumukh Aithal, Szymon Antoniak, Tarun Kumar Vangani, Teven Le Scao, Théo Cachet, Thibaut Lavril, Thomas Chabal, Thomas Coste, Thomas Defard, Thomas Foubert, Thomas Robert, Thomas Wang, Tianyu Zhang, Tim Lawson, Timothée Lacroix, Tobias Kronlachner, Tom Edwards, Tomas Hodan, Tuhin Das, Tyler Wang, Ulrick BLE, Umar Jamil, Umberto Tomasini, Valentin Macé, Van Phung, Vedant Nanda, Victor Jouault, Victor Letzelter, Victor Paltz, Victor Poucheret, Vincent Maladière, Vincent Pfister, Virgile Richard, Vladislav Bataev, Wen Ding Li, William Havard, William Marshall, Xinghui Li, Xingran Guo, Xinyu Yang, Yann Dreze, Yassine El Ouahidi, Yassir Bendou, Yihan Wang, Yves Martin des Taillades, Zaccharie Ramzi, Zhenlin Xu, Zsofia Csakany

arXiv:2607.25857v2cs.CLcs.CV

TL;DR

Existing safety classifiers rely on fixed taxonomies that do not accommodate heterogeneous moderation policies. Shieldstral reframes moderation as binary question answering and consolidates diverse data to train a 3B-parameter adaptive classifier. It matches or outperforms much larger models on text benchmarks and achieves state-of-the-art multimodal safety classification.

  • Problem

    Existing guardrail models use fixed safety taxonomies, while public safety datasets vary in category definitions and structure.

  • Method

    Shieldstral formulates moderation as binary question answering over natural-language safety queries and text or images, trained on curated heterogeneous data.

  • Results

    Shieldstral matches or outperforms nearly 7×-larger models across text benchmarks and achieves state-of-the-art multimodal safety classification.

  • Takeaways & Limitations

    The results support unified policy-adaptive moderation in which a small model can match or outperform much larger fixed-taxonomy models.

  • Takeaways & Limitations

    The adaptability evaluation relies on training and evaluation taxonomies that are independently designed and differ in structure, granularity, and category definitions.

Abstract

from arXiv · show

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

1 Introduction

Shieldstral is a 3B-parameter multimodal safety classifier that unifies content moderation as binary question answering over natural-language policies and text or images. With 54.1M curated samples, it matches or outperforms much larger models on text safety, achieves state-of-the-art multimodal safety, and supports policy-adaptive classification.

  • Guardrail models are essential for filtering harmful, biased, or illegal content as foundation models are deployed across diverse real-world applications.
  • Shieldstral reduces diverse moderation tasks to binary question answering by evaluating text and/or images against a natural-language safety query.The model produces a single continuous safety score rather than fixed categories.
  • 54.1M samples of careful data curation enable Shieldstral’s small 3B adaptive model to match or outperform much larger models.
  • 84.9% average F1 ranks Shieldstral at the top overall across diverse text safety benchmarks, matching or outperforming models nearly 7× its size.
  • 83.8% average F1 establishes state-of-the-art multimodal safety performance, while 91.3% F1 demonstrates policy-adaptive classification on a fine-grained taxonomy evaluation.Operators define moderation criteria through free-form natural-language queries at inference time.

2 Task Definition

Shieldstral reduces policy-adaptive safety classification to binary question answering, using structured instructions, yes/no queries, and multimodal documents. It trains with vocabulary-level cross-entropy and classifies using yes/no probabilities thresholded at 0.5.

  • Task formulation: Shieldstral structures safety classification as a standard binary question-answering task to support policy adaptation across diverse datasets.This reduction establishes a unified task formulation for the model.
  • Input structure: Each input contains a fixed system message and a user message with tagged instruction, query, and document fields.The instruction describes evaluation context and strictness, while the query asks a specific yes/no question.
  • Input structure: The document field can contain a user prompt, model response, formatted prompt–response pair, or image optionally accompanied by text.This structure accommodates both text and multimodal safety classification inputs.
  • Inference: At inference, Shieldstral compares yes and no token logprobs, computes a softmax-normalised safety score, and thresholds it at τ=0.5.Training uses standard cross-entropy loss over the full vocabulary at the output position.

3 Training Data Construction

Shieldstral’s training-data strategy unifies heterogeneous safety datasets into a common yes/no instruction–query–document format and scales them to approximately 54.1M samples. Contrastive curation and generation improve policy-specific discrimination by pairing content with relevant and irrelevant safety queries.

  • Data scale and coverage: 54.1M samples combine 45.2M open-source text, 4.4M synthetic contrastive text, and 4.5M multimodal samples across heterogeneous safety domains.Sources span safety, toxicity, hate speech, jailbreak detection, content moderation, and response quality.
  • Unified representation: A template-based instruction–query–document format converts prompt classification, response moderation, refusal detection, and toxicity detection into a unified yes/no task.Dataset-specific processors define labeling logic, category mappings, and tailored instruction templates, while paraphrase variants diversify phrasing.
  • Contrastive curation: Contrastive curation pairs identical content with matching and non-matching queries, forcing category discrimination rather than a coarse safe-versus-unsafe split.Positive queries can be binary, category-specific, or target-group-specific; negatives include absent categories, unrelated groups, and safe-content pairings.
  • Quality control: An open-source LLM filters samples whose dataset labels disagree with its binary or per-category classifications, reducing label noise across heterogeneous sources.The filtering removes both harmful-labeled benign samples and benign-labeled harmful samples.
  • Contrastive generation: Contrastive generation rewrites safe texts into unsafe variants for a target category while avoiding a sibling category, producing positive and negative queries.The LLM receives the safe text, target category, and sibling category, then generates the unsafe rewrite and both query types.

4 Adaptability Evaluation

The adaptability evaluation uses an independently designed taxonomy and contrastive, fixed-query dataset to test discrimination under policy drift and novel categories. Its separation from training labels, structures, and generation processes prevents performance from being explained by taxonomy memorization.

  • Motivation and evaluation design: Adaptability requires datasets covering policies that drift from training categories and entirely novel categories absent from training.The evaluation therefore extends contrastive generation to construct an adaptability-focused dataset.
  • Taxonomy design: The evaluation taxonomy is independently designed with disjoint categories, type-based distinctions, action-oriented names, and at least two sibling categories per subcategory.These principles ensure each content piece maps to one leaf while supporting sibling-category discrimination.
  • Taxonomy design: The taxonomy contains 12 super classes, 26 subcategories, and 52 leaf categories, evaluated with 90 manually authored canonical queries applied uniformly to all samples.Fixed queries decouple evaluation from template randomness and isolate the model’s adaptability.
  • Contrastive evaluation: For each category and sibling, contrastive generation produces a positive target-matching sample and a negative sibling-matching sample under an iso-query setting.Both samples receive the same target-category query, requiring discrimination of the specific harm type rather than generic unsafety.
  • Taxonomy separation: Training and evaluation taxonomies differ in structure, granularity, category definitions, names, and groupings, preventing strong evaluation performance from being attributed to memorized training labels.The training taxonomy has 4 severity tiers, whereas the evaluation taxonomy has 3 severity tiers and exactly 2 leaves per subcategory.

5 Model Architecture

Shieldstral is built on Ministral-3B-Base-2512, a 3B-parameter causal language model with native multimodal support through a Pixtral vision encoder.

  • Shieldstral uses Ministral-3B-Base-2512, a 3B-parameter Mistral-3-family causal language model with native multimodal support via Pixtral.

6 Training Recipe

Shieldstral is trained efficiently with LoRA in two complementary data regimes: public-only training improves benchmark calibration, while public-plus-generated training supports fine-grained, policy-adaptive generalisation. The final checkpoint combines these specialised models with the base instruct model through weighted pairwise SLERP merges.

  • Training: LoRA fine-tunes language-model parameters with cross-entropy on the single output token, matching full SFT without significant difference while improving training efficiency.Two specialised checkpoints are trained: P on public safety data and PG on public plus generated taxonomy data.
  • Model merging: P is well-calibrated to standard benchmarks, whereas PG provides fine-grained category discrimination but may suffer from distribution drift.The two regimes therefore offer complementary strengths without requiring additional training to combine them.
  • Model merging: 0.6 PG, 0.3 P, and 0.1 Ministral-3B-Instruct are the SLERP weights for policy-adaptive generalisation, benchmark calibration, and instruction-following capability, respectively.PG combines public and generated taxonomy data; P uses public data only; Ministral-3B-Instruct is the base instruct checkpoint.
  • Model merging: Pairwise SLERP merges produce the final checkpoint, with each component’s effect evaluated through ablation in Section 7.4.SLERP denotes spherical linear interpolation merging.

7 Evaluation

Shieldstral is evaluated across 16 benchmarks and 21 splits against 10 baselines, where its 3B parameters match larger models on text safety and achieve state-of-the-art multimodal performance. Ablations show policy-adaptive generalisation from curated and generated data, with LoRA selected for efficiency and SLERP merging improving the final model.

  • Evaluation setup: 16 benchmarks, 21 splits, and 10 baselines comprise the held-out evaluation used to assess Shieldstral fairly.All evaluation samples are held out from training data.
  • Text safety evaluation: 84.9% overall F1 lets the 3B Shieldstral match the 20B GPT-OSS-Safeguard-20B despite being the smallest compared model.Shieldstral also performs strongly on prompt classification, response classification, multilingual evaluation, and refusal detection.
  • Policy adaptability: 94.1% F1 is achieved by GPT-OSS-Safeguard-20B on adaptability, while Shieldstral’s final fine-grained taxonomy result reaches 91.3% F1 after SLERP merging.GPT-OSS-Safeguard-20B benefits from per-category policy prompts, reasoning, and its 20B parameter count.
  • Multimodal evaluation: 83.8% overall F1 makes Shieldstral the top multimodal model, ahead of OmniGuard’s 77.6%, while leading on two of three benchmarks.Shieldstral leads on VLGuard and UnsafeBench; LlavaGuard-7B leads its namesake benchmark at 81.4%.
  • Generalisation to unseen policies: 61.1% F1 from public-data fine-tuning rises by +23.3% to 84.4% with generated taxonomy data, despite evaluation categories being independently designed.The taxonomy uses different category names, granularity, and groupings, supporting generalisation rather than memorisation.

8 Conclusion

Shieldstral is a 3B-parameter policy-adaptive safety classifier that frames content moderation as binary question answering. Its training pipeline unifies heterogeneous sources through template-based unification and contrastive sample generation, supporting text and multimodal safety classification.

  • Fine-grained taxonomy validation: 86.1 Merge Acc., 92.0 Prec., 85.7 Rec., and 88.7 F1 are reported for 0.6PG+0.3P+0.1I.
  • Shieldstral is a 3B-parameter policy-adaptive safety classifier that formulates content moderation as a binary question-answering task.
  • Template-based unification and contrastive sample generation consolidate heterogeneous sources within a carefully designed training and data-curation pipeline.

A Multilingual Evaluation Results · B Full Evaluation Taxonomy

Shieldstral performs strongly overall in multilingual evaluation but underperforms in Prompt Classification for Arabic, Indonesian, and other low-resource languages. The evaluation taxonomy comprises 12 super classes, 26 subcategories, and 52 leaf categories spanning diverse safety domains.

  • A Multilingual Evaluation Results: Shieldstral performs strongly overall but underperforms in Prompt Classification for Arabic, Indonesian, and other low-resource languages.The limitation is specifically reported for multilingual evaluation results.
  • B Full Evaluation Taxonomy: The evaluation taxonomy contains 12 super classes, 26 subcategories, and 52 leaf categories.Table 9 presents the complete taxonomy structure.
  • B Full Evaluation Taxonomy: The taxonomy covers physical, financial, identity, system, and account crimes, including theft, fraud, identity deception, hacking, malware, phishing, and account takeover.Examples include Physical Property, Financial Crime, Identity Crime, System Attacks, and Account Attacks.
  • B Full Evaluation Taxonomy: It also includes personal, confidential, and self-harm risks, such as PII disclosure, doxxing, trade secrets, suicide promotion, and health risks.These categories appear under Personal Data Exposure, Confidential Data, and Self Harm.
  • B Full Evaluation Taxonomy: Additional classes address child safety, manipulation, reputation harm, and election integrity through categories such as child abuse, emotional blackmail, defamation, and voter suppression.The taxonomy lists paired subcategories within each of these domains.
  • B Full Evaluation Taxonomy: The remaining domains cover state security, media and commercial theft, ecosystem damage, animal harm, and drug crimes.Listed examples include espionage, terrorism, piracy, plagiarism, technology theft, pollution, animal cruelty, poaching, drug distribution, and drug manufacturing.

C Training vs. Evaluation Taxonomy Comparison

Shieldstral uses a broad, heterogeneous training taxonomy and a compact, symmetric evaluation taxonomy with distinct query strategies. The taxonomies remain aligned at a high level while refining or excluding selected training categories for evaluation.

  • Structural differences: The training taxonomy consolidates heterogeneous source taxonomies into 11 super classes and 73 leaf categories with uneven granularity.SC3 contains 15 categories across 5 subcategories, whereas SC5 contains 2 categories.
  • Structural differences: The evaluation taxonomy contains 12 symmetric super classes, 52 leaf categories, and exactly two leaf categories per subcategory for contrastive hard-negative evaluation.Each super class has 2–3 subcategories, with strict disjointness between sibling categories.
  • Query strategy: Training uses 110 prompt- and response-based templates sampled randomly, while an LLM rewriter generates a fresh query for each rewritten sample.Each of the 11 training super classes has five prompt-based and five response-based templates.
  • Query strategy: Evaluation assigns one fixed prompt query and one fixed response query to every leaf, subcategory, and super class, enabling deterministic reproducibility.The evaluation set includes 52 leaf categories, 26 subcategories, and 12 super classes.
  • Shared super-class alignment: Ten of 12 evaluation super classes align directly with training classes, but evaluation splits Criminal Activity into three classes and excludes System Security & Manipulation.The excluded training class covers jailbreak and prompt injection, which dedicated external benchmarks already evaluate.

D Complete Category-Level Results · E LLM Prompts for Taxonomy Data Generation

The appendices provide complete fine-grained taxonomy results and specify three LLM prompting strategies for generating training and evaluation data. The prompts enforce category-specific rewrites, matched positive/negative evaluation pairs, or coarse-grained unsafe rewrites.

  • D Complete Category-Level Results: Table 12 reports F1 scores (%) across 12 super classes, 26 subcategories, and 52 leaf categories for all evaluated models.Results are organized by hierarchy level, with the best score per row highlighted.
  • D Complete Category-Level Results: Drug Crimes results include scores for Drug Distribution and Drug Manufacturing, with unreliable one-sample scores omitted and strict/loose mappings averaged where marked.The table also notes a 0.5 threshold and reasoning_effort=high for applicable results.
  • E LLM Prompts for Taxonomy Data Generation: Taxonomy data is generated by rewriting safe content into unsafe variants using three prompting strategies selected by generation context.The strategies cover category-specific training rewrites, dual-version evaluation rewrites, and binary rewriting.
  • E.1 Category-Specific Rewriting (Training): For training data, each safe sample is rewritten into a target category while avoiding a specified sibling category, and the LLM generates positive and negative queries.The category-specific prompt applies to standalone text and prompt–response pairs.
  • E.1 Category-Specific Rewriting (Training): Training outputs require rewritten text in the original language plus diverse English yes/no queries about the target and negative categories.The required fields are REWRITTEN_TEXT, POSITIVE_QUERY, and NEGATIVE_QUERY, with varied phrasing, synonyms, and perspectives.
  • E.1 Category-Specific Rewriting (Training): The response variant prepends the original prompt as context and instructs the LLM to rewrite only the response.This adapts category-specific rewriting to prompt–response data.
  • E.2 Dual-Version Rewriting (Evaluation): For evaluation, one API call creates positive and negative rewrites of the same source text sharing one query, producing matched pairs for contrastive evaluation.The positive version targets the target category while excluding the negative category; the negative version covers only the negative category and excludes target-category terms.
  • E.3 Binary Rewriting: Binary rewriting makes safe text unsafe without a specified category, allowing any natural safety violation and generating a diverse safety-related yes/no query.All LLM calls use a safety-data-generation system message and temperature 0.7.
Loading 2607.25857v2…