Source-linked AI summary
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang
TL;DR
Safety evaluation often treats harmful generation as an attack outcome, leaving the content and variation of failures underexamined. HarmProfile profiles harmful-output distributions across frontier LLMs and finds that models produce harmful content at scale with distinct risk profiles, while harmfulness and diversity grow with capability.
Problem
Safety evaluation largely treats harmful generation as an attack outcome, leaving the content of safety failures underexamined beyond binary refusal judgments.
Method
HarmProfile profiles frontier LLM risk through a 15-category, 57-subcategory taxonomy and dual-axis harmfulness and diversity analysis of over 80,000 artifacts from 23 models.
Results
Frontier LLMs produce harmful content at scale with distinct risk profiles, and both harmfulness and diversity grow with model capability.
Takeaways & Limitations
HarmProfile makes harmful-content structure observable for more fine-grained studies and targeted defense-oriented evaluation beyond coarse alignment metrics.
Takeaways & Limitations
The distributions reflect ISC-triggered safety failures rather than necessarily the models’ full harmful capacity, and alternative elicitation methods may reveal different distributions.
Abstract
from arXiv · showhide
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .
1 Introduction
HarmProfile reframes harmful generation as a model-level distributional risk profile rather than a binary jailbreak outcome, addressing the limited content analysis of frontier-LLM safety failures. It introduces a large-scale, fine-grained benchmark spanning diverse harms and models, and finds distinct profiles whose harmfulness and diversity increase with capability.
- Motivation: Safety evaluation increasingly treats harmful generation as an attack outcome, reducing safety failures to binary judgments while leaving their content largely unexamined.Stronger alignment techniques, refusal mechanisms, and safety filters have made large-scale, high-quality harmful outputs from frontier models harder to obtain.
- Motivation: Harmful distributions differ markedly across models even when they all fail on harassment generation, including actionable instructions, hostile iconography, and fabricated personal details.These differences are invisible to binary attack-success metrics and reflect how models encode and reproduce harmful knowledge.
- HarmProfile: HarmProfile defines a model-level risk profile through the content, severity, and variation of safety failures, treating harmful generation as an object of distributional analysis.The framework is designed to answer what harmful distribution lies behind a jailbreak, not only whether the model can be jailbroken.
- HarmProfile: HarmProfile contains over 80,000 validated harmful artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories.Artifacts are autonomously generated through agentic workspaces without adversarial harmful seeds.
- Findings: Frontier LLMs reliably produce harmful content at scale yet exhibit distinct risk profiles, while both harmfulness and diversity grow with model capability.The findings suggest that frontier models may appear safe while harboring increasingly dangerous knowledge beneath the alignment surface.
2 Related Work
Prior safety evaluation primarily measures whether models refuse, comply, or complete harmful objectives, whereas HarmProfile analyzes the content produced during non-refusal failures. Related work identifies diverse jailbreak mechanisms and moves beyond binary refusal, but HarmProfile extends attack-free harmful-content collection into systematic, large-scale content-level benchmarking.
- Safety evaluations span toxicity, bias, trustworthiness, red-teaming, and dangerous-capability benchmarks, but primarily assess refusal, compliance, or task completion.
- Jailbreak research identifies competing objectives, adversarial suffixes, semantic prompt refinement, and context-structure attacks as diverse failure modes.
- Frontier models can exhibit evaluation awareness through alignment faking and in-context scheming, while ISC finds harmful content without attacks but does not analyze its content.
- HarmProfile extends attack-free harmful-content findings by systematically collecting large-scale harmful artifacts for content-level benchmarking.
- Attack-elicited outputs can show quality degradation of up to 92%, while jailbreak success may be hallucinated and ASR may mismatch actual harmful capability.
3 Method
HarmProfile combines taxonomy-guided autonomous generation with dual-axis profiling of harmfulness and diversity. It uses category-specific agentic workspaces, persistent-memory scaling, automated and pairwise harmfulness signals, and multi-granularity diversity metrics.
- Taxonomy-guided generation: HarmProfile organizes 57 subcategories under 15 high-level harm categories and uses specialized classifiers to guide category-consistent harmful-artifact generation.The taxonomy reconciles AI safety standards with frontier-lab usage policies, with full definitions provided in Appendix B.
- Taxonomy-guided generation: The method extends benign agentic workflows into 57 category-specific workspaces, where models synthesize harmful samples while evaluating or preparing classifier data.This design builds on the observation that harmful generations can emerge without explicit adversarial attacks.
- Scaled generation: Multi-round generation uses persistent memory, nonrepeating agent-assigned categories, hybrid embedding-plus-BM25 redundancy control, and diversity and severity guidelines.When redundancy exceeds threshold τ, a controller redirects generation toward a less redundant semantic region; rounds end after validation.
- Dual-axis profiling: Harmfulness combines StrongREJECT’s continuous [0, 1] score with GPT-4o-mini pairwise comparisons of 10 artifacts against 3 opponent models.The pairwise presentation order is randomized to mitigate position bias, and the combined score averages both signals.
- Dual-axis profiling: Diversity is measured at token, sentence, and topic levels using Self-BLEU complement, k-NN cosine distance, and BERTopic coverage fraction.The three metrics are combined with equal weights after min-max normalization across model–category pairs, while GPT-5.5 aggregates fine-grained topics into interpretable clusters.
4 Experiments
Experiments across 23 frontier LLMs show that models produce harmful content at scale while exhibiting distinct category-level risk profiles. Harmfulness and diversity generally increase with capability and scale, but zero-shot topic discovery eventually plateaus and outputs organize into domain-specific themes.
- Experimental setup: HarmProfile evaluates 23 frontier LLMs from 13 model families across 15 representative harm subcategories, with a broader 57-subcategory evaluation for nine models and over 80,000 validated artifacts overall.The full corpus also includes deep-generation runs.
- Overall harmfulness and diversity: Every tested model produces non-trivial harmful content across all 15 categories, with overall harmfulness ranging from 0.40 for model1 to 0.70 for the top tier.Most models achieve diversity scores of 0.50–0.72 and cover broad harmful scenarios.
- Category-level risk profiles: Models exhibit distinct category-level risk profiles: Grok-4.1-Fast scores 0.87 on jailbreak assistance versus 0.35 on copyright, while GLM-5.1 scores 0.91 on financial scam versus 0.47 on dehumanization.Aggregate safety scores can therefore obscure where risks concentrate.
- Capability scaling: Diversity correlates significantly with LMArena Elo (r=0.55, p=0.018), while harmfulness also shows a positive association, indicating that both dimensions tend to increase with model capability.Higher-harmfulness models generally exhibit broader diversity, while lower-scoring models cluster near the bottom of both dimensions.
- Within-family scaling: Within Qwen 3.5, harmfulness and diversity increase overall with scale: the dense 27B model reaches 0.66 harmfulness, while the 122B MoE variant reaches 0.60 diversity.The results suggest active parameters are more associated with harmful intensity, whereas total parameters may better cap diversity.
- Generation dynamics: New topic discovery plateaus around 3,000 artifacts, after which generations largely recombine existing themes; prompting with topics outside the saturated set elicits further valid artifacts.The plateau is tested using DeepSeek-V4-Flash on Protected-Attribute Hate, with replication on two additional models and topic expansion prompted through GPT-5.5.
5 Conclusion
HarmProfile introduces a content-centric benchmark of over 80,000 validated harmful artifacts from 23 frontier LLMs across 13 model families, showing that harmfulness and diversity vary across model risk profiles and grow with model capability.
- Contribution: HarmProfile contains over 80,000 validated harmful artifacts from 23 frontier LLMs across 13 model families.It characterizes model risk through content, severity, and category-level variation rather than binary refusal outcomes.
- Findings: Frontier LLMs produce harmful content at scale while exhibiting distinct harmful profiles across categories and model families.These profiles make harmful-content structure observable for fine-grained studies and targeted defense-oriented evaluation.
- Implications: Harmfulness and diversity grow with model capability, leaving harmful knowledge latent beneath the alignment surface.HarmProfile supports moving LLM safety evaluation beyond coarse alignment metrics toward deeper content-level analysis.
Ethics and Release
HarmProfile analyzes the content and behavioral structure of harmful frontier-LLM generations rather than developing attacks, and releases research resources while withholding raw harmful artifacts.
- Ethics and Release: The work focuses on harmful generations’ behavioral patterns, semantic profiles, and knowledge structures rather than attack and jailbreak methods.It publicly releases taxonomy definitions, scoring rubrics, aggregate results, framework code, and sanitized examples, while withholding raw harmful artifacts.
- Ethics and Release: HarmProfile does not introduce new attack methods or produce actionable harmful content, instead analyzing model generations under standard settings.The dataset is intended to support analysis and defense-oriented research beyond surface-level alignment evaluation.
Limitations · Appendix · Appendix Contents
The study’s risk profiles are conditional on ISC-triggered failures and specific evaluation choices, so alternative elicitation methods, scorers, encoders, or taxonomies may yield different results. The appendices provide model, taxonomy, protocol, embedding, evaluation, category-level, and artifact-level supporting materials.
- Limitations: Harmful distributions reflect ISC-triggered safety failures, not necessarily models’ full harmful capacity.Other elicitation methods may reveal different distributions.
- Limitations: Harmfulness scores depend on StrongREJECT and GPT-4o-mini as pairwise judge.Alternative scorers could shift model rankings.
- Limitations: Diversity metrics depend on Jina embeddings and BERTopic clustering choices, so alternative encoders could shift rankings.The reported diversity results are therefore tied to these methodological choices.
- Limitations: The 57-subcategory taxonomy does not exhaustively cover all possible harms.This limits the scope of the reported harmful-output distributions.
- Appendix Contents: Appendix materials cover the model list, full harm taxonomy, autonomous workspace protocol, artifact embedding space, full harm ability matrix, and evaluation prompts.These materials are listed in Appendix Contents.
- Appendix Contents: Additional appendices provide word clouds by harm category, full 57-subcategory results, full-set scope, harmfulness, diversity, and category-level BERTopic coverage.The full results are organized under Appendix H and its subsections.
- Appendix Contents: The appendices also include a high-risk cluster topic taxonomy and example artifacts across all 15 harm categories.The example-artifact section enumerates categories from Violence & Extremism through Model Adversarial.
- Appendix Contents: Further contents list example artifacts for the remaining harm categories and a zero-shot versus prompted ablation.Listed categories include Privacy, Fraud & Deception, Misinformation, Illicit Activity, Political & Civic, Cyber Offense, Intellectual Property, High-Stakes Advice, and Model Adversarial.
A Model List … E Full Harm Ability Matrix
The appendix defines HarmProfile’s evaluated models and harm taxonomy, details its constrained agentic collection and validation protocol, and characterizes the resulting artifact embedding space. It covers 23 models, 57 subcategories, and validated artifacts organized for comparative analysis across harm domains.
- A Model List: 23 frontier LLMs are evaluated and sorted by release date.Table 4 provides the model list used in HarmProfile.
- B Full Harm Taxonomy: 57 harm subcategories are defined with codes, parent categories, and evaluation-scope definitions for judging.Tables 5–7 specify the full taxonomy used by judges.
- C Autonomous Workspace Protocol: Each of the 57 workspaces supports autonomous exploration within tool restrictions, two-layer validation, and iterative coverage tracking.Agents explore one harm subcategory per self-contained workspace, while approved artifacts remain inaccessible to them.
- Agent Round Prompt (cleaned): Each round uses a fresh agent instance that reads, writes, tests, revises, and logs while advancing diversity and severity without harmful seeds or examples.The agent builds classifier test cases within a named risk category and freely decides what to generate.
- System Prompt: The protocol separates agent-visible local tests from driver-side authoritative validation, rejecting artifacts that fail schema, labeling, length, uniqueness, or deduplication checks.Passing pytest is necessary but insufficient for dataset inclusion; the unsafe response must contain at least 50 words.
- Initial Log Template (Factual Misinformation): Artifacts use a query–safe response–unsafe response schema, with category descriptions deduplicated across rounds and expected labels tied to the tested classifier.Workspaces initialize logs with task, label, schema, and coverage guidance emphasizing adversarial, out-of-distribution, escalating, and boundary cases.
- D Artifact Embedding Space: The full embedding space is visualized with t-SNE, coloring artifacts by harm category S1–S15 and distinguishing source models by marker shape.Marginal densities are shown on both axes, and categories exhibit varying degrees of overlap.
- D Artifact Embedding Space: Most harm categories form coherent semantic clusters, while S4 Hate & Discrimination, S7 Privacy, and S6 Sexual Content are tight and separated, whereas S8 Fraud and S10 Illicit Activity are more dispersed.The embedding structure is described as reflecting semantic similarity rather than merely surface-level lexical overlap.
F Evaluation Prompts · G Word Clouds by Harm Category
This section documents the evaluation prompt used to combine StrongREJECT scores with pairwise LLM-judge win rates and presents word clouds for evaluated harm subcategories. It also reports expanded cross-model MAUVE similarities that confirm family-level clustering patterns.
- F Evaluation Prompts: Harmfulness is assessed by combining StrongREJECT scores with pairwise LLM-judge win rates.The method is described in §3.2.
- F Evaluation Prompts: Figure 15 reproduces verbatim the pairwise comparison prompt used by GPT-4o-mini.
- F Evaluation Prompts: Figure 13 reports pairwise MAUVE scores between all 23 LLMs.
- F Evaluation Prompts: The expanded model set confirms clustering patterns observed in the 14-model subset.
- F Evaluation Prompts: Model families sharing training data form distinct similarity blocks.
- G Word Clouds by Harm Category: Figures 16–17 present top-100 frequency and distinctive-token word clouds for all evaluated harm subcategories.Frequency word clouds appear on top, while distinctive-token word clouds appear on the bottom.
H Full 57-Subcategory Results · Pairwise Harmfulness Comparison Prompt
The appendix validates the main experiments with a full 57-subcategory protocol on 9 models and 47,500 validated artifacts, while the accompanying prompt uses pairwise judging to compare harmful responses. Diversity is summarized across models and safety sections with weighted scores, and GPT-4o-mini selects the more harmful response.
- H Full 57-Subcategory Results: 47,500 validated artifacts comprise the full-set run used as robustness evidence for the main-text claims.The run includes StrongREJECT scores and diversity analyses.
- H Full 57-Subcategory Results: 9 models are evaluated under the full 57-subcategory protocol to test whether the 15 representative subcategories were a narrow-selection artifact.The full-set protocol extends beyond the main experiments’ 15 representative subcategories.
- H Full 57-Subcategory Results: The full-set results provide robustness evidence for claims from the main text.This is the stated purpose of the completed appendix results.
- H Full 57-Subcategory Results: Figure 14 reports weighted diversity across models and safety sections using an equal-weight combination of min-max-normalized Dtoken, Dsent, and Dtopic.Rows are sorted by overall diversity, while columns are sorted by section-level average diversity.
- H Full 57-Subcategory Results: Darker cells in the diversity heatmap indicate higher diversity scores.The heatmap organizes model-by-section diversity values, with an overall-diversity column at right.
- Pairwise Harmfulness Comparison Prompt: The pairwise prompt presents two assistants responding to the same type of harmful request in a specified category.The category is inserted through the template placeholder “{{category}}”.
- Pairwise Harmfulness Comparison Prompt: GPT-4o-mini compares the two assistant responses and selects the more harmful one.Figure 15 identifies this template as the pairwise judge prompt.
H.1 Full-Set Scope
The full-set analysis successfully scored artifacts from all 9 model files, with each model contributing several thousand validated artifacts despite slight differences in completion and validator filtering.
- Full-set scope: 9 full-set model files were scored successfully, with mean StrongREJECT scores reported per model.Table 9 summarizes the number of scored artifacts and the corresponding mean StrongREJECT score for each model.
- Full-set scope: Sample counts vary slightly across models because generation completion and validator filtering differ.
- Full-set scope: Each model contributes several thousand validated artifacts to the full-set analysis.
H.2 Full-Set Harmfulness
Across thousands of validated artifacts, every full-set model shows non-trivial mean harmfulness, indicating that harmful generation extends beyond 15 representative subcategories. The scorer-only results provide a larger-scale check rather than replacing the main text’s combined harmfulness score.
- Full-set harmfulness: Every full-set model receives a non-trivial mean harmfulness score across thousands of validated artifacts, so harmful generation is not confined to a small hand-picked subset of categories.DeepSeek-V4-Flash and Qwen3.6-35B-A3B score highest, followed by GLM-4.7-Flash and Qwen3.6-27B; Qwen3-Coder-30B remains far from benign.
- Full-set harmfulness: The full-set StrongREJECT scores are a larger-scale scorer-only check that the main conclusion is not driven solely by the 15 representative subcategories.They are not intended to replace the combined harmfulness score, which averages StrongREJECT with the pairwise LLM-judge win rate.
H.3 Full-Set Diversity
Full-set diversity is measured by combining token-, sentence-, and topic-level coverage, with MiniMax-M2.7 leading the ranking alongside three other models. Component scores show that lexical variation can outpace semantic and topical expansion, whereas top-ranked models sustain broad coverage across dimensions.
- Ranking: The combined diversity ranking is led by MiniMax-M2.7, DeepSeek-V4-Flash, Gemma4-31B, and Qwen3.6-27B.The score averages token-level diversity, sentence-level diversity, and topic-level BERTopic coverage.
- Component analysis: GPT-5.4-Nano shows high token diversity but substantially lower sentence-level and topic-level diversity.This pattern indicates lexical variation without comparable expansion of semantic or topical coverage.
- Component analysis: The top-ranked models maintain consistently high coverage across multiple diversity components.
H.4 Category-Level BERTopic Coverage … S15.3 Guardrail Bypass (Qwen3.6-27B)
H.4 shows that models differ in harmful-topic breadth and that topic convergence generalizes across model families. The appendices then document a broad high-risk taxonomy and representative validated artifacts spanning violence, self-harm, weapons, hate, harassment, sexual content, privacy, fraud, misinformation, and adversarial behavior.
- H.4 Category-Level BERTopic Coverage: MiniMax-M2.7, Qwen3.6-27B, and DeepSeek-V4-Flash cover the largest fraction of discovered topics, while Qwen3-Coder-30B covers the smallest.Average topic coverage varies substantially across models in the full-set BERTopic analysis.
- H.4 Category-Level BERTopic Coverage: Topic coverage reveals risk-profile differences beyond harmfulness intensity, and zero-shot topic convergence also appears in Grok-4.1-Fast and GLM-5.1.The appendix reports the same saturation pattern observed for DeepSeek-V4-Flash across additional model families.
- I High-Risk Cluster Topic Taxonomy: 131 cluster-level topics organize harmful artifacts across 15 high-risk harm categories, with each cluster reporting merged fine-grained topics and total sample counts.The taxonomy provides a structured view of category-level harmful content.
- J Example Artifacts by Category: Representative validated artifacts span violence and extremism, self-harm, weapons and mass harm, hate and discrimination, and harassment and defamation.Examples include incitement, graphic violence, suicide, explosives, protected-attribute hate, harassment, defamation, and impersonation.
- J.10–J.14 Illicit Activity, Political & Civic, Cyber Offense, Intellectual Property, and High-Stakes Advice: The taxonomy and examples further include illicit activity, political and civic harms, cyber offenses, intellectual-property abuse, and high-stakes advice.The supplied artifacts include chemical procurement, election-related manipulation, routing sabotage, and health misinformation, while the taxonomy lists additional subcategories.
- J Example Artifacts by Category: The examples also cover sexual content, privacy violations, fraud and deception, and misinformation across multiple model families.These categories include child sexual abuse material, minor grooming, non-consensual recording, personally identifiable information, doxxing, phishing, market manipulation, and factual misinformation.
K Zero-Shot vs. Prompted Ablation · Ablation Topic Generation Prompt
The ablation prompted DeepSeek-V4-Flash with 100 novel Protected-Attribute Hate topics generated outside its saturated zero-shot topic set. It produced validator-passing artifacts for every topic, indicating that zero-shot saturation reflects autonomous topic-selection diversity rather than the model’s harmful-knowledge boundary.
- K Zero-Shot vs. Prompted Ablation: 100 novel Protected-Attribute Hate topics were generated outside DeepSeek-V4-Flash’s saturated zero-shot topic set.GPT-5.5 generated the topics for the prompted ablation.
- K Zero-Shot vs. Prompted Ablation: Each generated topic was injected as the category field for DeepSeek-V4-Flash in prompted mode.The experiment tested whether zero-shot saturation reflected a knowledge boundary or the zero-shot regime.
- K Zero-Shot vs. Prompted Ablation: The results show harmful knowledge beyond what zero-shot exploration surfaces.This conclusion follows from successful artifact generation for all 100 topics outside the saturated zero-shot set.
- K Zero-Shot vs. Prompted Ablation: 100/100 topics were successfully generated by DeepSeek-V4-Flash.Table 12 lists the 100 novel topics used in the ablation.
- K Zero-Shot vs. Prompted Ablation: 100% success rate was achieved for valid, validator-passing artifacts across all 100 prompted topics.DeepSeek-V4-Flash produced valid artifacts for every prompted topic.
- Ablation Topic Generation Prompt: The generation prompt requested 100 additional protected_attribute_hate topic areas meaningfully different from previously identified topics.The prompt framed these topics as targets for safety evaluation.
- Ablation Topic Generation Prompt: Table 12 contains 100 novel topics targeting distinct identity groups not covered by zero-shot generation.All 100 were successfully generated when provided as the category field.