Source-linked AI summary

On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Yue Huang, Chujie Gao, Siyuan Wu, Haoran Wang, Xiangqi Wang, Yujun Zhou, Yanbo Wang, Jiayi Ye, Jiawen Shi, Qihui Zhang, Yuan Li, Han Bao, Zhaoyi Liu, Tianrui Guan, Dongping Chen, Ruoxi Chen, Kehan Guo, Andy Zou, Bryan Hooi Kuen-Yew, Caiming Xiong, Elias Stengel-Eskin, Hongyang Zhang, Hongzhi Yin, Huan Zhang, Huaxiu Yao, Jaehong Yoon, Jieyu Zhang, Kai Shu, Kaijie Zhu, Ranjay Krishna, Swabha Swayamdipta, Taiwei Shi, Weijia Shi, Xiang Li, Yiwei Li, Yuexing Hao, Zhihao Jia, Zhize Li, Xiuying Chen, Zhengzhong Tu, Xiyang Hu, Tianyi Zhou, Jieyu Zhao, Lichao Sun, Furong Huang, Or Cohen Sasson, Prasanna Sattigeri, Anka Reuel, Max Lamparth, Yue Zhao, Nouha Dziri, Yu Su, Huan Sun, Heng Ji, Chaowei Xiao, Mohit Bansal, Nitesh V. Chawla, Jian Pei, Jianfeng Gao, Michael Backes, Philip S. Yu, Neil Zhenqiang Gong, Pin-Yu Chen, Bo Li, Dawn Song, Xiangliang Zhang

arXiv:2502.14296v5cs.CY

TL;DR

GenFMs’ expanding use raises the need to define and assess trustworthiness across safety, fairness, privacy, robustness, ethics, and advanced AI risks. The paper proposes multidisciplinary guidelines and TrustGen, a dynamic benchmark for three model categories, finding substantial progress alongside persistent gaps and emphasizing continued evaluation and collaboration.

  • Problem

    GenFMs’ widespread adoption and vulnerabilities create an unresolved need for a unified framework defining and assessing trustworthiness across dimensions and contexts.

  • Method

    The paper synthesizes governance and industry practices into guidelines and introduces TrustGen, a dynamic evaluation framework for text-to-image, language, and vision-language models.

  • Results

    Most evaluated GenFMs show substantial trustworthiness progress, while open-source models rapidly close the gap with proprietary models and critical gaps remain.

  • Takeaways & Limitations

    Responsible GenFM integration requires ongoing collaboration, rigorous evaluation, and continuous refinement of models and trustworthiness frameworks.

  • Takeaways & Limitations

    Strengthening safety and alignment without preserving utility can overly constrain models and reduce useful or creative responses.

Abstract

from arXiv · show

Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions. This paper presents a comprehensive framework to address these challenges through three key contributions. First, we systematically review global AI governance laws and policies from governments and regulatory bodies, as well as industry practices and standards. Based on this analysis, we propose a set of guiding principles for GenFMs, developed through extensive multidisciplinary collaboration that integrates technical, ethical, legal, and societal perspectives. Second, we introduce TrustGen, the first dynamic benchmarking platform designed to evaluate trustworthiness across multiple dimensions and model types, including text-to-image, large language, and vision-language models. TrustGen leverages modular components--metadata curation, test case generation, and contextual variation--to enable adaptive and iterative assessments, overcoming the limitations of static evaluation methods. Using TrustGen, we reveal significant progress in trustworthiness while identifying persistent challenges. Finally, we provide an in-depth discussion of the challenges and future directions for trustworthy GenFMs, which reveals the complex, evolving nature of trustworthiness, highlighting the nuanced trade-offs between utility and trustworthiness, and consideration for various downstream applications, identifying persistent challenges and providing a strategic roadmap for future research. This work establishes a holistic framework for advancing trustworthiness in GenAI, paving the way for safer and more responsible integration of GenFMs into critical applications. To facilitate advancement in the community, we release the toolkit for dynamic evaluation.

1 Introduction

GenFMs’ broad capabilities and social impact make trustworthy behavior difficult to define and assess across diverse contexts. The paper responds with unified guidelines, TrustGen’s dynamic evaluation, and discussion of persistent challenges and future directions.

  • GenFMs support diverse applications but face risks including jailbreaks, safety-filter bypasses, privacy leaks, and unpredictable or unethical behavior.Their broad generalization across tasks and contexts complicates consistent trustworthiness assessment.
  • Opaque architectures, massive scale, heterogeneous training data, and continual updates complicate interpretability, accountability, safety consistency, and traceability.These properties require evaluation across diverse tasks and contexts.
  • Static large-scale evaluations are unsustainable because new models and changing user needs require repeated dataset, metric, and methodology design.The process is time-consuming and inflexible, motivating adaptive assessment.
  • TrustGen evaluates text-to-image, large language, and vision-language models using overall trustworthiness scores summarized in Figures 4–6.The benchmark is presented as comprehensive and adaptive.
  • Most evaluated models achieve relatively high trustworthiness scores, yet significant bottlenecks remain and high scores do not guarantee reliability in every context.Open-source models can also match or surpass proprietary models, with CogView-3-Plus outperforming DALL-E-3.
  • The paper concludes that trustworthiness requires interdisciplinary discussion of evaluation, technical strategies, societal effects, downstream implications, and future research.These directions aim to align GenFMs with human values and societal expectations.

2 Background

The background reviews corporate trust-and-safety practices and existing evaluation benchmarks to identify requirements and gaps for unified guidelines and adaptive assessment. TrustGen is positioned as a broad benchmark spanning primary trustworthiness aspects and multiple GenFM types.

  • 2.1 Approaches to Enhancing Trustworthiness From Corporate: Corporate trustworthiness practices are examined to identify essential features of trustworthy GenFMs and inform unified guidelines.The review covers prominent developers and their approaches to real-world trust.
  • 2.2 Evaluation of Generative Models: Existing evaluation methods and benchmarks provide strengths but leave gaps that motivate a more adaptive and effective assessment benchmark.The review explicitly informs TrustGen’s benchmark development.
  • 2.1 Approaches to Enhancing Trustworthiness From Corporate: OpenAI combines red teaming, model system cards, safety standards, alignment work, secure infrastructure, AI-generated-content identifiers, and democratic-input programs.Its stated principles include minimizing harm, building trust, iteration, and broadly distributed benefits.
  • 2.1 Approaches to Enhancing Trustworthiness From Corporate: Meta uses pre-deployment stress tests, multilingual moderation, prompt-attack detection, cybersecurity benchmarks, and responsible deployment partnerships for Llama models.These measures address malicious use, unsafe content, prompt injection, jailbreaking, and cybersecurity risks.
  • 2.1 Approaches to Enhancing Trustworthiness From Corporate: Microsoft emphasizes equitable AI, bias mitigation, transparency, social-good initiatives, trustworthy-AI principles, and privacy and security commitments.Its initiatives span healthcare, conservation, data interpretation, environmental challenges, and government applications.
  • 2.3 Trustworthiness-Related Benchmark: TrustGen covers truthfulness, safety, fairness, robustness, privacy, machine ethics, and advanced AI risk across text-to-image, language, and vision-language models.Its data-construction strategies and modules support dynamic and diverse testing.

3 Guidelines of Trustworthy Generative Foundation Models

The paper proposes flexible, stakeholder-adaptive guidelines for GenFMs because trustworthiness is multifaceted and context-dependent. The framework integrates legal, ethical, risk-management, user-centered, and adaptability considerations into principles covering core trustworthiness dimensions.

  • Guideline Motivation: Trustworthiness varies by application context, so the guidelines avoid rigid universal rules and support diverse stakeholders.They are intended for developers, regulators, organizations, and researchers.
  • Guideline Motivation: The guidelines are application-agnostic and stakeholder-adaptive, distinguishing them from broad policy-oriented frameworks such as the EU AI Act and AI Bill of Rights.Their purpose is to provide a flexible foundation for GenFM-specific decisions and evaluations.
  • Guideline Content: The principles require fairness, transparency, human oversight, accountability, robustness, harmlessness, reliable information, uncertainty communication, and privacy protection.These requirements span model development, deployment, user interactions, adversarial conditions, and data protection.
  • Guideline Content: Guideline 7 requires reliable and accurate information while requiring models to communicate uncertainty when information is uncertain or speculative.The paper notes that absolute accuracy is difficult because of data, training, and output-measurement limitations.
  • Guideline Content: The framework emphasizes privacy and data protection for both user-provided information and information generated about users during interactions.This extends privacy considerations beyond only the initial input.
  • Guideline Development: Guideline development considers legal compliance, ethics and social responsibility, risk management, user-centered design, and adaptability.The process drew on corporate and governmental principles, policies, and regulations through multidisciplinary collaboration.

4 Designing TrustGen, a Dynamic Benchmark Platform for Evaluating the Trustworthiness of GenFMs

TrustGen is an open-source dynamic benchmark for evaluating trustworthiness across three GenFM categories and seven dimensions. Its modular, continuously updated pipeline addresses the obsolescence, memorization, and scalability problems of static benchmarks while combining automation with human validation.

  • Platform Overview: TrustGen evaluates text-to-image, large language, and vision-language models across seven trustworthiness dimensions using broad metrics.Its design aims to provide thorough assessments across model types and dimensions.
  • Motivation: Static benchmarks become outdated, can be memorized by models, and are impractical to repeatedly rebuild as models and user needs evolve.These limitations motivate a shift toward dynamic evaluation.
  • Platform Architecture: TrustGen uses Metadata Curator, Test Case Builder, and Contextual Variator modules to keep datasets and evaluations continuously updated.The modules respectively curate metadata, generate test cases, and vary contexts and question formats.
  • Design Principles: TrustGen balances trustworthy behavior with practical utility, prioritizes realistic low-cost attacks, and supports reproducible open-source construction.The toolkit enables users to create evaluation datasets and replicate the benchmark process.
  • Platform Architecture: The Metadata Curator processes existing data or retrieves information, while the Test Case Builder creates labeled evaluation cases from generative or programmatic operations.The construction process preserves corresponding ground-truth labels for generated inputs.
  • Platform Architecture: The Contextual Variator paraphrases and reformats cases to address prompt sensitivity and limited diversity in template-based generation.Human evaluators check for semantic shifts and acceptable data quality after variation.

5 Benchmarking Text-to-Image Models

TrustGen’s T2I evaluations show progress alongside persistent weaknesses in truthfulness, safety, fairness, robustness, and privacy. Models generally perform well on several dimensions, but complex scenes, perturbations, and differing privacy-content types expose important gaps.

  • Truthfulness: All mainstream T2I models underperform in truthfulness, with Dall-E 3 achieving the highest score.Dall-E 3 incorporates more entities and attributes than the evaluated open-source models.
  • Truthfulness: T2I models struggle with complex prompts containing multiple objects, spatial relationships, and global scene attributes.Human annotators observed strong aesthetics and stylistic coherence alongside difficulty organizing complex scenes.
  • Safety: Dall-E 3 achieves the highest Safety Score at 94, while SD-3.5-large and SD-3.5-large-turbo record the lowest scores at 47 and 53, respectively.The paper associates Dall-E 3’s result with its external moderation system and the lower scores with weaker filtering or greater prompt sensitivity.
  • Fairness: HunyuanDiT leads fairness with 95.5, while SD-3.5-large records the lowest score at 91.83.Overall fairness scores are relatively close, but the results still indicate varying performance across models.
  • Robustness: Robustness scores range from 92.98 to 94.77 after perturbation, with Playground-v2.5 lowest and Kolors highest.The results indicate slight instability compared with clean inputs.
  • Privacy: Privacy performance differs by content type: Dall-E 3 scores 59.38 for organization-related privacy and 72.22 for individual-related privacy.The paper identifies this discrepancy as evidence that filtering is more effective for personal information than organizational data.

6 Benchmarking Large Language Models

TrustGen’s LLM assessment finds progress across trustworthiness dimensions, but substantial and uneven weaknesses remain in hallucination, sycophancy, honesty, safety, privacy, robustness, and fairness-related behaviors.

  • Hallucination: Most LLMs perform better on dynamically generated datasets than on established benchmarks, especially for question answering, with some fact-checking exceptions.The exceptions include Llama-3.1-8B and Llama-3.1-70B.
  • Sycophancy: Sycophancy varies sharply across models: persona-induced accuracy changes range from 1.30% for o1-preview to 100% for Qwen-2.5-72B.For preconception sycophancy, Gemini-1.5-pro changes 1.01% versus 37.92% for GPT-3.5-turbo.
  • Sycophancy: Models often change truthful answers after doubtful follow-ups, with QwQ-32B changing 19.19% of responses versus over 88% for several Gemini and Claude models.This pattern is described as self-doubt sycophancy in multi-round dialogue.
  • Honesty: Honesty remains limited and imbalanced: leading models stay below 75%, while SIC honesty can reach zero even as most LIES scores exceed 80%.The results indicate a need for more diverse training samples in categories where honesty is lacking.
  • Jailbreak, fairness, privacy, and robustness: Trustworthiness performance is uneven across other dimensions: proprietary models lead jailbreak resistance, stereotype accuracy does not ensure disparagement performance, and smaller models often preserve privacy better.Open-weight and proprietary models show similar toxicity levels, while robustness leaders reach 99.36%.
  • Safety: Most LLMs show low overall toxicity and strong exaggerated-safety performance, but some models remain over-cautious and extreme toxic cases persist.Most models have less than 5% full RtA and under 10% combined RtA for exaggerated safety.

7 Benchmarking Vision-Language Models

TrustGen evaluates vision-language models across hallucination, jailbreak, fairness, privacy, ethics, and robustness dimensions. Results show substantial variation: models can perform strongly on some tasks while remaining vulnerable in specific contexts.

  • 7.2 Hallucination: GPT-4o and Claude-3.5-Sonnet achieve the highest overall accuracy across both hallucination benchmarks.
  • 7.2 Hallucination: 17.91% separates top-performing models from lower-performing models in hallucination-inducing scenarios.The comparison is between GPT-4o or Claude-3.5-Sonnet and models such as Claude-3-Haiku and Llama models.
  • 7.3 Jailbreak: 99.9% is Claude-3.5-Sonnet’s average refuse-to-answer rate across five jailbreak attacks.Only the FigStep attack succeeds against this model; open-source models show lower refusal rates.
  • 7.5 Robustness: Joint image-text perturbations cause the most substantial performance degradation across the VQA robustness settings.Image-only perturbations have minimal impact, while perturbations can produce both positive and negative effects.
  • 7.6 Privacy: 93.81% is achieved by the smaller Llama-3.2-11B-V on privacy, exceeding Qwen-2-VL-72B at 51.37%.The comparison indicates that model scale alone does not determine privacy performance.

8 Other Generative Models

Trustworthiness concerns extend beyond text, image, and vision-language models to video, audio, agents, and other multimodal systems. Existing work develops benchmarks and mitigations, but reported vulnerabilities remain across these model classes.

  • 8.1 Any-to-Any Models: Any-to-any models create a critical safety gap because comprehensive investigations of their safety implications remain limited.The concern covers systems that combine multiple modalities, including language, vision, and other outputs.
  • 8.2 Video Generative Models: 12 safety aspects are covered by T2VSafetyBench for text-to-video model assessment.The benchmark uses malicious prompts created with LLMs and jailbreak attacks.
  • 8.3 Audio Generative Models: Audio generative models raise deepfake risks because high-fidelity systems can imitate individuals’ voices without consent.The paper connects this risk to fraudulent activities such as impersonation scams and unauthorized access.
  • 8.3 Audio Generative Models: Fairness, robustness, and privacy remain critical audio-model issues alongside misuse through audio deepfakes.Non-diverse training data can favor certain accents or dialects and underperform for others.
  • 8.4 Generative Agents: Larger model parameters do not guarantee better agent performance because training data and response strategies also affect tool utilization.
  • 8.4 Generative Agents: None of 16 tested agents surpass a 60% safety score in Agent-SafetyBench.The benchmark spans 349 interaction environments, 2,000 test cases, and eight safety-risk categories.

9 Trustworthiness in Downstream Applications

Downstream deployment makes trustworthiness application-specific: medical, embodied, transportation, image-generation, and legal systems face distinct safety, privacy, fairness, accuracy, and security requirements. The paper presents promising responses but emphasizes that substantial challenges remain.

  • 9.1 Medicine: Medical GenFM deployment requires reliability, fairness, privacy, and interpretability because errors can affect clinical decisions and patient groups.Bias from underrepresented demographics can produce inaccurate or unfair outcomes.
  • 9.3 Embodied AI: Embodied systems require safeguards, behavior validation, anomaly detection, and misuse prevention as autonomy increases.
  • 9.4 Autonomous Driving: System-wide GenFM transportation deployments introduce security, privacy, safety, and robustness challenges tied to sensitive traveler and infrastructure data.
  • 9.7 Image Generation: Image memorization is more frequent in text-to-image models trained on small-to-medium datasets than in models trained on larger, diverse datasets.ImageNet-trained models show minimal or undetectable replication in the cited comparison.
  • 9.8 Law: 58% is the minimum hallucination frequency reported for LLMs in legal work.Other AI-powered legal tools are reported to hallucinate between 17% and 33% of the time during legal analysis.
  • 9.8 Law: TrustGen’s truthfulness, fairness, and privacy focus is especially relevant for assessing GenFM suitability in legal settings.

10 Further Discussion

Trustworthiness in generative models is dynamic, stakeholder-dependent, and intertwined with utility, safety, alignment, and ethical reasoning. The discussion highlights adaptive evaluation, balanced safeguards, and persistent variation across models and contexts.

  • Dynamic Trustworthiness: Dynamic trustworthiness requires adaptive evaluation because static metrics may miss context-specific demands across domains and stakeholders.Transparent benchmark assumptions help stakeholders interpret which evaluations fit their needs.
  • Dynamic Trustworthiness: Trustworthiness is not fixed; deployment and assessment must adapt to changing needs across domains.The paper characterizes trustworthiness as a complex, multidimensional quality that requires continual renegotiation.
  • Trustworthiness and Utility: Overemphasizing safety can constrain useful or creative responses, showing that trustworthiness and utility require careful balancing.The discussion identifies stringent filtering and rigid ethical frameworks as possible sources of reduced utility.
  • Trustworthiness and Utility: Utility without fairness, transparency, or resistance to manipulation can produce harmful outputs and is unsustainable in high-stakes settings.Healthcare and finance are cited as environments where untrustworthy models are unlikely to be adopted sustainably.
  • Trustworthiness and Utility: A harmlessness-first approach treats safety as a foundation for improving utility, while the paper argues that both goals should be pursued together.The discussion frames trustworthiness as a necessary component of utility rather than merely a constraint on it.
  • Ethical Reasoning: Models differ in ethical behavior, ranging from neutrality to decisive or emotionally driven choices, and their decisions may not capture real-world moral complexity.The discussion also contrasts top-down clarity with bottom-up flexibility and inconsistency, motivating interdisciplinary research and reflective equilibrium.

11 Conclusion

The paper presents a holistic framework for defining and evaluating GenFM trustworthiness, combining multidisciplinary guidelines with TrustGen’s dynamic assessments. Results show substantial progress alongside critical gaps, including open-source models rapidly narrowing the trustworthiness gap with proprietary models.

  • The framework covers safety, fairness, privacy, robustness, machine ethics, and advanced AI risks.
  • TrustGen enables continuous, flexible trustworthiness assessments across text-to-image, large language, and vision-language models.
  • The evaluation finds substantial trustworthiness advances while uncovering critical gaps requiring rigorous, ongoing oversight.
  • Open-source models are rapidly closing the trustworthiness gap with proprietary models.
  • The paper identifies persistent ethical, legal, and societal challenges as priorities for future trustworthy-GenFM research.

Diversity Statement

The research brings together expertise from many technical, social-scientific, medical, legal, and applied domains to study trustworthy generative models.

  • The project integrates perspectives from NLP, computer vision, human-computer interaction, computer security, medicine, and computational social science.
  • It also includes expertise in robotics, data mining, law, and AI for science.
  • Each field contributes distinct perspectives to the research on trustworthy generative models.

Contribution Statement

This section documents student contributions across trustworthiness-related sections and lists the generative models included in the paper’s model introductions. The models span language, multimodal, and text-to-image systems from open-source and commercial developers.

  • Student contributors’ involvement is organized across sections using trustworthiness-dimension colors, with leaders marked by stars.
  • The model descriptions cover capabilities such as multimodal processing, long-context support, multilingual dialogue, reasoning, and efficient deployment.
  • GLM-4V-Plus and Qwen2-VL-72B represent vision-language models designed for visual reasoning and multimodal processing.
  • Playground 2.5 and CogView-3-Plus represent text-to-image models using diffusion-based or DiT-based architectures.

B.1.2 Robustness Image Description Generation

The robustness image-description procedure generates structured descriptions for text-to-image models and applies several perturbations to test response stability. Perturbations include social tagging, multilingual blending, distractive text, syntactic disruptions, recondite words, and topic extraction.

  • The procedure first generates an original image description from a supplied aspect in compact JSON format.
  • Social tagging and recondite words define perturbation categories for modifying the input description.
  • Multilingual blending translates the sentence into Chinese and requests the result without explanation.
  • Distractive text introduces a brief off-topic digression, while syntactic disruptions add common grammatical mistakes.
  • The recondite-words prompt replaces 1–4 common words with rarer synonyms while preserving direct output.

C.1 Jailbreak Results of Large Language Models

This section presents detailed RtA results under jailbreak attacks and examples spanning harmful-content categories, honesty, persuasion, privacy, and organizational information privacy.

  • Jailbreak evaluation: Table 30 reports refuse-to-answer (RtA) results under different kinds of jailbreak attacks.The supplied caption identifies the evaluation measure and attack variation but provides no numerical values.
  • Safety categories: The examples include sexual content, gore and violence, hate, drugs and illegal activities, and horror.These examples are presented as NSFW image-description instances.
  • Honesty: The section includes criteria and examples for six dimensions used to construct honest LLMs.The supplied passage identifies the table’s scope but does not enumerate the six dimensions.
  • Privacy and persuasion: The section also provides instances of persuasion strategies and guides to personal and organizational information-privacy aspects.The supplied figure descriptions identify privacy aspects for web retrieval and organizational information privacy.

D.5 Data Examples For LLM Fairness

The examples show fairness-related failures across language and vision-language models, including stereotype judgments, spatial and visual errors, harmful advice, and successful jailbreaks.

  • LLM fairness: Gemma-2-27B labels a poverty-and-eviction statement as a stereotype, while Claude-3.5-Sonnet incorrectly identifies a power-outage statement as a stereotype.The examples distinguish the models’ selected answers from the provided correctness labels.
  • Visual reasoning: GPT-4V and LLaVA produce conflicting answers on state geography, with supplied labels marking several Delaware answers wrong and one GPT-4V answer correct.The examples include Arizona, Oklahoma, North Carolina, and Delaware comparisons.
  • Visual reasoning: GPT-4V incorrectly denies a streetlight, while another example contrasts GPT-4V’s spacesuit description with Claude’s statue interpretation.These cases illustrate object-recognition and image-interpretation disagreements.
  • Safety and social reasoning: The VLM examples include an incorrectly handled Crimea image pair, rural-versus-urban interpretation, reputation-harm advice, and successful jailbreak outputs.The jailbreak examples involve finding damaging information, public shaming, and amplifying religious tensions.

E Proof: Indirect Generation Mitigates VLM Interior Bias

The proof compares direct question generation from images with indirect generation through compressed descriptions, arguing that indirect generation has lower domain-knowledge contamination under stated entropy assumptions.

  • Generation processes: The indirect process generates q_indirect through a description d=g(i), whereas the direct process generates q_direct directly from image i.The description is treated as a compressed representation of the image.
  • Main inequality: The proof derives I(K;q_direct|i) > I(K;q_indirect|d).The result compares contamination-related mutual information for direct image generation and indirect description-based generation.
  • Assumptions: The proof assumes H(q_direct|i) > H(q_indirect|d) and H(q_direct|K,i) < H(q_indirect|K,d).These are introduced as hypotheses describing conditional uncertainty in the direct and indirect processes.
  • Information-theoretic formulation: Conditional mutual information measures the dependence between domain knowledge K and generated questions given the input.K denotes the VLM’s domain-knowledge space, including prior knowledge, biases, and latent representations.
  • Implication: Because contamination is proportional to the corresponding mutual information, the indirect method reduces contamination from domain knowledge K and mitigates bias in VLM outputs.The argument defines contamination for the direct process and concludes the analogous comparison for the indirect process.
Loading 2502.14296v5…