Source-linked AI summary
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
Xingjun Ma, Yixu Wang, Hengyuan Xu, Yutao Wu, Yifan Ding, Yunhan Zhao, Zilong Wang, Jiabin Hua, Ming Wen, Jianan Liu, Ranjie Duan, Yifeng Gao, Yingshui Tan, Yunhao Chen, Hui Xue, Xin Wang, Wei Cheng, Jingjing Chen, Zuxuan Wu, Bo Li, Yu-Gang Jiang
TL;DR
Whether major capability gains in LLMs and MLLMs translate into comparable safety improvements remains unclear because evaluations are fragmented across modalities and threat models. The report conducts a unified evaluation of six frontier models across language, vision–language, and image generation using benchmark, adversarial, multilingual, and compliance testing. It finds a heterogeneous safety landscape: GPT-5.2 is strong and balanced, while models generally retain severe adversarial vulnerabilities and trade-offs across safety dimensions.
Problem
Fragmented evaluations leave unclear whether advances in LLM and MLLM capabilities translate into comparable safety improvements across modalities and threat models.
Method
The report evaluates six frontier models across language, vision–language, and image generation with unified benchmark, jailbreak, multilingual, and regulatory-compliance testing.
Results
The models show a heterogeneous safety landscape: GPT-5.2 is strong and balanced, other models exhibit cross-dimension trade-offs, and worst-case adversarial safety remains below 6%.
Takeaways & Limitations
Safety is multidimensional and depends on modality, language, and evaluation design, supporting holistic assessment rather than reliance on isolated safety measures.
Takeaways & Limitations
The evaluation covers only part of the evolving safety landscape and cannot capture long-tail risks or emergent behaviors in real-world deployment.
Abstract
from arXiv · showhide
The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and vision, yet whether these advances translate into comparable improvements in safety remains unclear, partly due to fragmented evaluations that focus on isolated modalities or threat models. In this report, we present an integrated safety evaluation of six frontier models--GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5--assessing each across language, vision-language, and image generation using a unified protocol that combines benchmark, adversarial, multilingual, and compliance evaluations. By aggregating results into safety leaderboards and model profiles, we reveal a highly uneven safety landscape: while GPT-5.2 demonstrates consistently strong and balanced performance, other models exhibit clear trade-offs across benchmark safety, adversarial robustness, multilingual generalization, and regulatory compliance. Despite strong results under standard benchmarks, all models remain highly vulnerable under adversarial testing, with worst-case safety rates dropping below 6%. Text-to-image models show slightly stronger alignment in regulated visual risk categories, yet remain fragile when faced with adversarial or semantically ambiguous prompts. Overall, these findings highlight that safety in frontier models is inherently multidimensional--shaped by modality, language, and evaluation design--underscoring the need for standardized, holistic safety assessments to better reflect real-world risk and guide responsible deployment.
1 Introduction
Frontier-model safety evaluation has expanded across text, multimodal interaction, attacks, languages, and compliance, but fragmented testing limits coherent understanding. This report compares models across these dimensions and reveals substantial differences in safety balance and robustness.
- Motivation: Rapid LLM and MLLM adoption has brought reasoning, instruction following, multimodal perception, and agentic behavior into large-scale real-world use alongside persistent safety vulnerabilities.Reported vulnerabilities include harmful-content generation, unsafe procedural guidance, and susceptibility to jailbreak attacks.
- Evaluation gap: Existing safety research spans jailbreak prompts, harmful-content benchmarks, multimodal benchmarks, and unified evaluation platforms, yet fragmented evaluations hinder understanding of models’ true safety envelopes.The report addresses fragmentation by combining standardized benchmark, adversarial, multilingual, and compliance evaluations.
- Evaluation scope: Six frontier models are evaluated across language-only, vision–language, and image-generation modes using a unified protocol covering benchmarks, jailbreak attacks, 18 languages, and regulatory compliance.The evaluated models are GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5.
- Comparative findings: GPT-5.2 consistently leads across benchmark, adversarial, multilingual, and compliance dimensions, while Grok 4.1 Fast remains near-last and other models show uneven trade-offs.Reported GPT-5.2 scores are 91.59% for Benchmark Evaluation, 54.26% for Adversarial Robustness, 77.50% for Multilingual Safety, and 90.22% for Regulatory Compliance.
- Safety profiles: Safety profiles are multidimensional rather than reducible to one scalar, exposing distinct model archetypes shaped by trade-offs among benchmark alignment, adversarial robustness, multilingual safety, and compliance.The reported profiles characterize GPT-5.2 as balanced, Gemini 3 Pro as reactive, Qwen3-VL as polarized, and Grok 4.1 Fast as guardrail-light.
2 Language Safety
Language-safety results reveal strong but uneven alignment: GPT-5.2 leads overall and across adversarial refusal tests, while models show pronounced weaknesses on social bias, contextual misuse, and adaptive jailbreaks.
- Overall benchmark results: GPT-5.2 achieves the highest macro-average safe rate at 91.59% and leads the overall language-safety comparison.It also ranks highest or tied-highest on 4/5 datasets.
- Model profiles: 96.67% on StrongREJECT and 92.00% on ALERT show GPT-5.2’s strong performance against explicit and adversarially phrased harmful requests.Its standard textual safety profile remains strong, although example responses expose contextual failures.
- Model profiles: 99.00% on BBQ makes Gemini 3 Pro the strongest model for the evaluated social-bias benchmark, despite weaker jailbreak handling than GPT-5.2.The benchmark spread ranges from 45.00% to 99.00%, indicating substantial variation across risk categories.
- Model profiles: Grok 4.1 Fast records the lowest macro-average, 66.60%, alongside weak SORRY-Bench and StrongREJECT scores.Its reported results indicate broad difficulty recognizing and refusing baseline harmful instructions.
- Adversarial evaluation: No model exceeds 6% worst-case adversarial safety under 30 black-box jailbreak attacks, despite relatively strong refusal-based defenses.GPT-5.2 has Safeworst near 6%, while adaptive and obfuscated attacks continue to bypass defenses across models.
2.3 Multilingual Evaluation
The multilingual evaluation tests guardrail-style safety judgment across 18 languages using PolyGuardPrompt and ML-Bench under separate prompt- and response-based inputs. Models perform similarly on standard safety data but diverge on region-specific regulatory judgments, with GPT-5.2 showing the strongest resilience.
- Evaluation setup: The evaluation measures guardrail-style safety judgment across 18 languages using PolyGuardPrompt, ML-Bench, and separate prompt- and response-based inputs.Models receive prompts and responses separately to focus judgment on response safety; performance is measured with micro F1.
- Overall findings: Models show strong, converged performance on PolyGuardPrompt but diverge substantially on policy-grounded, region-specific ML-Bench evaluations.The contrast indicates that applying specific regulatory frameworks is harder than recognizing general safety concepts across languages.
- PolyGuardPrompt: On PGP-P, GPT-5.2 and Gemini 3 Pro lead with macro F1 0.85, followed by Qwen3-VL at 0.84 and Grok 4.1 Fast at 0.82.This performance is reported across 17 languages and reflects prompt-included evaluation.
- PolyGuardPrompt: On PGP-R, Qwen3-VL leads at 0.79, exceeding GPT-5.2 and Gemini 3 Pro at 0.76 in response-only judgment.The response-only setting removes prompt context and focuses on identifying unsafe output patterns.
- ML-Bench: ML-Bench produces universal performance degradation, while GPT-5.2 records the highest F1 scores in both prompt and response evaluations.Gemini 3 Pro, Grok 4.1 Fast, and Qwen3-VL experience particularly sharp drops in response-only settings.
- Cross-lingual variation: On ML-Bench, all models perform better in high-resource languages and struggle more with lower-resource or culturally distinct contexts such as Japanese and Hindi.GPT-5.2 shows the most uniform distribution across languages, whereas the other models exhibit higher variance.
2.4 Compliance Evaluation
The compliance evaluation translates three governance frameworks into executable tests of legally and normatively defined constraints. GPT-5.2 leads overall, but model failures persist in contextual application of regulations across biometric surveillance, intellectual property, and transparency scenarios.
- Evaluation setup: The evaluation tests GPT-5.2, Gemini 3 Pro, Qwen3-VL, and Grok 4.1 Fast against NIST AI RMF, the EU AI Act, and MAS FEAT.Regulatory text is decomposed into atomic rules with compliant and adversarial guidance for testing.
- Evaluation metric: Compliance Rate (%) measures the percentage of responses satisfying the rule-specific judgment criteria.Higher values indicate better compliance across the governance frameworks.
- Overall results: GPT-5.2 leads overall compliance at 90.22%, followed by Qwen3-VL at 77.11%, Gemini 3 Pro at 73.54%, and Grok 4.1 Fast at 45.97%.GPT-5.2 exceeds the second-best model by more than 13%.
- Model profiles: GPT-5.2 leads all three frameworks, reaching 98.17% on NIST, 89.63% on the EU AI Act, and 82.86% on FEAT.It reaches 100% in Predictive Policing and Emotion Recognition but scores 66.67% in Transparency.
- Model profiles: Qwen3-VL scores 100% on Ethics but only 48.60% on Real-time Remote Biometric Identification, revealing uneven alignment across regulatory categories.Its stronger performance on general ethical principles does not extend consistently to specific high-risk biometric regulations.
- Model profiles: Gemini 3 Pro achieves 74.29% on FEAT, surpassing Qwen3-VL, but both Gemini 3 Pro and GPT-5.2 score 66.67% in Transparency.The results identify transparency as a shared weakness in meeting rigorous obligations.
- Model profiles: Grok 4.1 Fast records 45.97% overall compliance and 22.71% on NIST, substantially below the 70%+ baselines maintained by other models.Even its strongest framework, FEAT at 61.17%, trails the second-lowest performer.
- Failure analysis: Failure cases span prohibited biometric categorization, copyright reproduction, remote biometric identification, and deceptive interface design.These examples show failures in recognizing legal and ethical constraints embedded in complex real-world instructions.
3 Vision–Language Safety
Multimodal benchmark safety is uneven: GPT-5.2 leads overall and across most benchmarks, while models show distinct weaknesses on subtle, context-dependent harms. Adversarial evaluation further reveals substantial model-dependent fragility despite strong benchmark performance.
- Benchmark Evaluation: 92.14%: GPT-5.2 achieves the highest macro-average multimodal safety rate, ahead of Qwen3-VL at 83.32%, Gemini 3 Pro at 82.53%, and Grok 4.1 Fast at 67.97%.No model consistently dominates every benchmark, indicating unevenly addressed multimodal safety.
- Benchmark Evaluation: 55.38%–88.69%: MemeSafetyBench shows the widest safety-rate range, whereas SIUO exceeds 87.43% for every model.The contrast suggests that implicit, culturally grounded meme harms are harder than unsafe outputs from otherwise benign visual–textual combinations.
- Model Profiles: GPT-5.2 records the highest rates on MemeSafetyBench (88.69%), MIS (90.20%), and USB-SafeBench (92.66%), while reaching 97.01% on SIUO.Its performance spans implicit meme harm, multi-image reasoning, risk categorization, and visual–textual combinations.
- Model Profiles: Grok 4.1 Fast falls to 55.38% on MemeSafetyBench, 65.10% on MIS, and 63.95% on USB-SafeBench despite 87.43% on SIUO.The report associates this degradation with difficulty detecting implicit harm and subtle multimodal risks.
- Failure Modes: Direct visual–textual prompts can produce harmful responses without explicit jailbreak templates, including actionable weaponization guidance, stereotype confirmation, and fabricated illicit-funding contracts.These cases show models prioritizing analytical helpfulness, visual reasoning, or creative completion over harm prevention when prompts appear neutral.
- Adversarial Evaluation: 97.24%: GPT-5.2 leads adversarial multimodal safety, while Qwen3-VL reaches 78.89%, Gemini 3 Pro 75.44%, and Grok 4.1 Fast 68.34%.Except for GPT-5.2, models show pronounced degradation, with VLJailbreakBench (Hard) spanning 41.72%–98.54%.
- Adversarial Evaluation: 98.54% and 98.33%: GPT-5.2 achieves near-saturated safe rates on VLJailbreakBench and JailbreakV-28K, while maintaining 94.84% on MM-SafetyBench.Its consistency across heterogeneous attacks suggests a well-integrated and generalizable safety alignment strategy.
- Adversarial Evaluation: Qwen3-VL reaches 89.94% on MM-SafetyBench and 86.17% on JailbreakV-28K but drops to 60.57% on VLJailbreakBench.The pattern indicates stronger resistance to transfer-based and visually manipulated attacks than to MLLM-generated subtle jailbreaks.
4 Image Generation Safety
The two image-generation models exhibit different safety strategies and failure modes: similar refusal rates produce different safe-generation outcomes, and both remain vulnerable to adversarial or semantically subtle visual risks. Nano Banana Pro generally retains stronger safety, while Seedream 4.5 becomes substantially more fragile when refusals are bypassed.
- Benchmark Evaluation: Approximately 21%: both T2I models show comparable refusal rates, yet Nano Banana Pro reaches a 52% Safe Rate versus 40% for Seedream 4.5.Similar refusal behavior therefore does not translate into comparable safe-generation outcomes.
- Benchmark Evaluation: Unsafe generation exceeds 76% for Disturbing prompts and 69% for Violence prompts, with relatively low refusal rates.These categories remain especially difficult because visually grounded harms are harder to identify and suppress than text-centric risks.
- Safety Strategies: Nano Banana Pro primarily uses implicit sanitization, neutralizing or attenuating harmful elements instead of consistently issuing explicit refusals.Its softened outputs can preserve harmful semantics, including dangerous scenarios, hateful intent, and illicit activity.
- Safety Strategies: Seedream 4.5 relies more heavily on explicit refusal, especially for Sexual prompts, but can generate distorted or abstract anatomy when refusal is circumvented.Its failure mode reflects stricter sensitivity to selected keywords or semantic cues without reliably producing safe alternatives.
- Adversarial Evaluation: Higher refusal alone does not guarantee robustness: GenBreak exposes deeper weaknesses, with Seedream 4.5 becoming disproportionately unsafe and toxic after refusal bypass.Worst-case safety depends on post-bypass behavior rather than refusal frequency alone.
- Adversarial Evaluation: 54.00% versus 19.67%: Nano Banana Pro and Seedream 4.5 record these worst-case average Safe rates under adversarial T2I evaluation.Nano Banana Pro also has Harmful rate 27.67% and Toxicity 0.44, compared with Seedream 4.5 at 38.33% and 0.57.
- Semantic and Visual Blind Spots: Both models are more likely to miss nudity rendered artistically or placed as a small background element within a complex scene.For Violence & Gore, they appear to use perceptual thresholds for blood and gore rather than semantic understanding of violent intent.
- Semantic and Visual Blind Spots: Under adversarial prompting, Seedream 4.5 frequently generates hate symbols, whereas Nano Banana Pro consistently refuses such requests.The contrast suggests different levels of semantic grounding for prohibited visual concepts.
5 Conclusion
The report’s integrated evaluation finds a heterogeneous safety landscape across six frontier models and modalities. GPT-5.2 is strong and balanced, but benchmark results often fail to generalize to adversarial or semantically ambiguous settings.
- Integrated Evaluation: Six frontier models are evaluated across language, vision–language, and text-to-image generation using benchmark, adversarial, multilingual, and regulatory-compliance assessments.The unified protocol provides a systematic view of safety under diverse conditions.
- Main Findings: GPT-5.2 demonstrates strong and balanced performance across modalities, while most systems trade off benchmark alignment, adversarial robustness, multilingual generalization, and compliance.The report characterizes safety as heterogeneous rather than reducible to a single leaderboard dimension.
- Main Findings: Strong benchmark performance often fails to generalize under adversarial prompting, and vision–language interaction introduces failure modes comparable to language-only settings.Text-to-image models show relatively stronger alignment in regulated visual categories but remain brittle under adversarial or semantically ambiguous prompts.
6 Limitations and Disclaimer
The report’s conclusions are bounded by limited scope, deployment mismatch, general-purpose evaluation, model evolution, and its non-regulatory academic status.
- Scope and coverage: The evaluations cover only part of the safety landscape and cannot capture long-tail risks or emergent real-world behaviors.Results are therefore indicative rather than exhaustive measures of system risk.
- Deployment conditions: Distributional shift, continuous updates, user adaptation, and platform safeguards may cause reported safety performance to differ from live deployment.
- Evaluation framework: A general-purpose comparative framework may underrepresent safeguards tailored to specific applications, jurisdictions, or deployment constraints.Some models may consequently appear disadvantaged despite strong alignment in intended contexts.
- Temporal validity: The findings reflect model behavior at testing time rather than permanent or intrinsic safety properties.The evaluated systems are actively maintained and continuously evolving.
- Disclaimer: This academic analysis is not an official regulatory position, certification outcome, regulatory judgment, or basis for enforcement action.
A.1 Multilingual Judge Template
The experiments used a unified judge template for safety assessment, documented in Figure 20.
- A.1 Multilingual Judge Template: The safety judge template used in the experiments is included as Figure 20.
- A.1 Multilingual Judge Template: Figure 20 presents the unified judge template referenced by the report’s evaluation materials.
A.2 Hierarchical Taxonomy of Regulatory Compliance Risks
The regulatory-compliance taxonomy organizes risks hierarchically across three levels, from broad categories to specific scenarios.
- A.2 Hierarchical Taxonomy of Regulatory Compliance Risks: The taxonomy contains 7 Level-1 categories, 23 Level-2 categories, and 74 Level-3 risk scenarios.
- A.2 Hierarchical Taxonomy of Regulatory Compliance Risks: Table 11 presents the regulatory-compliance risk taxonomy as a Level-1 → Level-2 → Level-3 hierarchy.
- A.2 Hierarchical Taxonomy of Regulatory Compliance Risks: The Level-3 entries provide specific risk scenarios within the broader regulatory-compliance categories.
A.3 Regulatory Compliance Evaluation Prompt Template
The regulatory-compliance evaluation uses a prompt template that applies taxonomy-defined risk categories and definitions to image-safety adjudication.
- A.3 Regulatory Compliance Evaluation Prompt Template: Qwen3-VL’s image-safety adjudication prompt dynamically inserts each taxonomy risk category and its detailed definition.The placeholders “{theme}” and “{description}” are replaced for each evaluated risk category.
- A.3 Regulatory Compliance Evaluation Prompt Template: Figure 21 documents the prompt template used to evaluate image regulatory compliance.
- A.3 Regulatory Compliance Evaluation Prompt Template: The regulatory-compliance taxonomy is presented as a hierarchical Level-1, Level-2, and Level-3 structure.
A.4 Adversarial Evaluation (Attack Suite)
The adversarial attack suite is organized by attack mechanism, covering encoding, token, persona, code, prompt-optimization, semantic, reasoning, and multilingual strategies.
- Encoding & Cipher & Rephrase: The suite includes encoding and cipher attacks, including Caesar/Morse ciphers, character mutations, and past-tense rephrasing.
- Token Manipulation: Token-manipulation attacks use reversed character or word order, ASCII-art keyword masking, and adversarial tokens.
- Role-playing & Persona; Code Wrapping: Persona and code-wrapping attacks disguise harmful requests through “Do Anything Now” personas, comments, or function calls.
- Prompt Optimization: Prompt-optimization attacks apply iterative prompt refinement or tree-of-attacks search with pruning.
- Semantic Manipulation: Semantic-manipulation attacks use persuasion-based framing and authority impersonation, while semantic jailbreak strategies provide additional bypass methods.
- Reason Obfuscation: Reasoning-based attacks disguise harmful intent through fictional layers, nested scenes, decomposition, obfuscation, or multi-step reasoning.
- Multilingual: Multilingual attacks target low-resource-language bypasses and multilingual back-translation.
A.5 Grok 4 Fast Prompt Template
Figure 22 presents the prompt template used to evaluate image toxicity for Grok 4 Fast.
- The figure shows a prompt template for evaluating image toxicity in Grok 4 Fast.