Source-linked AI summary

AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap

Q. Vera Liao, Jennifer Wortman Vaughan

arXiv:2306.01941v2cs.HCcs.AIcs.CY

TL;DR

Powerful LLMs create risks, while transparency remains largely missing from current LLM discourse. This paper develops a human-centered roadmap by examining LLM-specific challenges, synthesizing HCI/FATE lessons, and assessing four transparency approaches, resulting in open questions and directions for future research.

  • Problem

    Transparency is largely missing from discourse around LLMs despite risks including bias, hallucinations, harmful content, and privacy or security threats.

  • Method

    The paper maps a human-centered transparency research roadmap by reflecting on LLM-specific challenges, synthesizing HCI/FATE lessons, and examining four transparency approaches.

  • Results

    The paper identifies open questions about applying model reporting, evaluation results, explanations, and uncertainty communication to LLMs.

  • Takeaways & Limitations

    Transparency approaches for LLMs should support human understanding and account for stakeholders’ different goals and contexts.

  • Takeaways & Limitations

    Proprietary LLMs often restrict access to weights, parameters, training-data details, and training resources, limiting transparency to black-box probing.

Abstract

from arXiv · show

The rise of powerful large language models (LLMs) brings about tremendous opportunities for innovation but also looming risks for individuals and society at large. We have reached a pivotal moment for ensuring that LLMs and LLM-infused applications are developed and deployed responsibly. However, a central pillar of responsible AI -- transparency -- is largely missing from the current discourse around LLMs. It is paramount to pursue new approaches to provide transparency for LLMs, and years of research at the intersection of AI and human-computer interaction (HCI) highlight that we must do so with a human-centered perspective: Transparency is fundamentally about supporting appropriate human understanding, and this understanding is sought by different stakeholders with different goals in different contexts. In this new era of LLMs, we must develop and design approaches to transparency by considering the needs of stakeholders in the emerging LLM ecosystem, the novel types of LLM-infused applications being built, and the new usage patterns and challenges around LLMs, all while building on lessons learned about how people process, interact with, and make use of information. We reflect on the unique challenges that arise in providing transparency for LLMs, along with lessons learned from HCI and responsible AI research that has taken a human-centered perspective on AI transparency. We then lay out four common approaches that the community has taken to achieve transparency -- model reporting, publishing evaluation results, providing explanations, and communicating uncertainty -- and call out open questions around how these approaches may or may not be applied to LLMs. We hope this provides a starting point for discussion and a useful roadmap for future research.

1 Introduction

LLMs offer broad transformative potential but introduce substantial risks, making responsible development and deployment urgent. The paper argues for a human-centered transparency roadmap tailored to stakeholders, applications, and usage contexts.

  • LLMs are being deployed across search, coding, productivity, and many industries, with expected effects on tasks and occupations.
  • LLMs can encode bias, hallucinate convincing inaccuracies, generate toxic content, and expose sensitive information.
  • Transparency supports appropriate understanding of models’ capabilities, limitations, operation, and output control, enabling responsible development and deployment.
  • Existing transparency approaches include model and dataset documentation, explanations of individual outputs, and uncertainty communication, but no single solution fits every context.
  • The paper maps a human-centered roadmap by examining LLM-specific challenges, HCI and FATE lessons, and four established transparency approaches.
  • The paper focuses on informational transparency while acknowledging normative, relational, and social dimensions, and notes that many issues extend to multimodal generative models.

2 What Makes Transparency for LLMs Challenging?

Transparency for LLMs is difficult because their capabilities, architectures, applications, stakeholders, and surrounding perceptions are complex and continually changing. Proprietary access, deployment pressures, and interacting system components further constrain what can be understood and disclosed.

  • Complex and Uncertain Model Capabilities and Behaviors: LLMs support diverse tasks and contexts, unlike models with narrowly defined inputs and outputs, while their behavior can be inconsistent, nondeterministic, and update-dependent.
  • Massive and Opaque Architectures: Their massive architectures and poorly documented training data prevent complete understanding of model knowledge, reasoning, and behavior.
  • Proprietary Technology: Most powerful LLMs are proprietary or API-accessible, limiting access to weights, parameters, training-data details, and internal workings.
  • New and Complex Applications: LLM-infused applications may combine multiple models, tools, plugins, and external services, so transparency must address their interactions rather than an isolated LLM.
  • Expanded and Diverse Stakeholders: The expanding ecosystem includes diverse developers, users, decision-makers, regulators, auditors, and impacted groups with different transparency needs.
  • Rapidly Evolving and Often Flawed Public Perception: Public perceptions of LLM capabilities and operation are evolving under the influence of media, marketing, events, and application design.
  • Organizational Pressure to Move Fast and Deploy at Scale: Pressure to release quickly and scale broadly can conflict with responsible AI efforts, requiring stronger governance, regulation, or organizational incentives.

3 What Lessons Can We Learn from Prior Research?

Prior HCI and responsible-AI research frames transparency as a human-centered means to support stakeholders’ goals, understanding, appropriate trust, and control. For LLMs, the roadmap emphasizes adapting communication and evaluation to new goals, dynamic mental models, and the risks that transparency itself can create.

  • 3.1 Transparency as a Means to Many Ends: A Goal-Oriented Perspective: Human-centered transparency should be evaluated by whether it helps stakeholders achieve their end goals, which vary by context and application stakes.Relevant goals include learning, decision-making, trust development, diagnosis, ideation, model adaptation, prompting, and discovering risky behaviors.
  • 3.1 Transparency as a Means to Many Ends: A Goal-Oriented Perspective: LLM transparency research should investigate new stakeholder goals, including ideation, model adaptation, prompting, and discovering risky model behaviors.Approaches should be developed and evaluated according to how well they support these goals.
  • 3.2 Transparency to Support Appropriate Levels of Trust: Transparency should support appropriate trust rather than trust indiscriminately, especially because explanations can increase overreliance on incorrect AI outputs.Trust may concern the base model, an LLM-infused application, its provider, or specific functions and outputs.
  • 3.3 Transparency and Control Often Go Hand-in-Hand: Transparency and control are intertwined design goals: understanding should be paired with mechanisms that let stakeholders act on, improve, or steer LLM behavior.The paper connects this agenda to interactive machine learning and calls for more participatory and inclusive control mechanisms.
  • 3.4 The Importance of Mental Models: Mental models are internal representations shaped by experience, and transparency must support both functional knowledge of use and structural knowledge of how and why systems work.Because mental models evolve through ongoing interactions, interpretability should be considered dynamically and situationally rather than as a single documentation or explanation intervention.
  • 3.5 How Information is Communicated Matters: Communication format matters for LLM transparency: natural-language systems can express uncertainty through hedging or refusal, while tailored explanations can improve error detection.Selective explanations adapted to recipients were reported as easier to process and more helpful for detecting model errors in an AI-assisted decision task.
  • 3.6 Limits of Transparency: Transparency can fail or cause harm when it produces information without understanding, shifts accountability burdens to users, or threatens privacy and security.Users lacking technical background may face higher burdens, and transparency can be misused to exploit trust or reliance.

4 What Existing Approaches Can We Draw On?

The paper organizes transparency for LLMs around model reporting, evaluation results, explanations, and uncertainty, while emphasizing human-centered goals and context. It highlights unresolved challenges involving documentation, evaluation scope, explanation faithfulness, and interaction design.

  • The paper identifies four transparency approaches: model and data reporting, evaluation results, explanations, and uncertainty communication.It also encourages new approaches, including model-interrogation tools and output-specific risk communication.
  • Model reporting: Model reporting for LLMs must address unclear input-output spaces, uncertain general-purpose capabilities, and incomplete training-data and development-background information.These categories are difficult because of LLM complexity, proprietary constraints, and uncertainty about capabilities.
  • Model reporting: Model reporting should extend beyond static cards to formats such as FAQs, onboarding pages, landing pages, and media communication.The paper notes that these formats can provide functional information and shape understanding, with standardization where appropriate.
  • Publishing evaluation results: Automated and human evaluations have limitations: ground truth may miss context-dependent quality, while human evaluation is costly and lacks standardized, reproducible practices.These concerns affect validity and generalizability to real-world settings.
  • Publishing evaluation results: Evaluation should account for its target stakeholder and purpose because practitioner needs may differ from NLP researchers’ goals.The authors argue that articulating distinct evaluation goals can support multiple coexisting techniques.
  • Publishing evaluation results: Behavioral evaluation should be participatory and iterative, yet risk taxonomies may lack use-case granularity and red-team practices remain insufficiently transparent.The paper calls for shared best practices to improve understanding of LLM risks.
  • Providing explanations: Explanations should be faithful, understandable, and useful, but LLM-generated explanations can contradict one another and diverge from the model’s actual process.GPT-4 may produce explanations that vary with precise inputs or conflict with its output.
  • Providing explanations: LLM explanations should be interactive, tailored, and compatible with evolving multi-turn dialogues rather than governed by one monolithic standard.The paper recommends distinguishing explanation types by mechanisms, stances, questions, contexts, limitations, and pitfalls.

5 Summary and Discussion

The paper consolidates a human-centered roadmap for LLM transparency, emphasizing diverse stakeholders, evolving models, and shared challenges across transparency approaches. It also identifies provenance, temporal change, external oversight, feedback mechanisms, and applicability beyond LLMs as priorities for future research.

  • The roadmap synthesizes LLM-specific challenges, HCI/FATE lessons, and four transparency approaches: model reporting, evaluation results, explanations, and uncertainty communication.It frames these approaches as requiring human-centered research rather than treating any single approach as sufficient.
  • Effective LLM transparency must account for base models, adapted models, applications, diverse stakeholders, and their surrounding socio-organizational contexts.The paper presents these interacting dimensions as common considerations across transparency approaches.
  • Transparency research should address AI-generated text provenance and make interaction with AI systems clear in relevant tasks.The paper connects this concern to regulatory discussions and proposed disclosure obligations.
  • Because base-model updates propagate into LLM-infused applications, research must track model provenance and communicate how changes affect end-users.Relevant provenance includes architecture, training, datasets, adaptation details, functions, and evaluation results.
  • External audits and contestation mechanisms remain open governance challenges, including selecting methods and metrics and addressing community-specific harms.The paper also calls for feedback systems that use reported failures to identify and address recurring patterns.
  • Many identified challenges and open problems also apply to large-scale generative models beyond LLMs, including multimodal models.The paper encourages additional transparency research as these models become more widespread.
Loading 2306.01941v2…