Source-linked AI summary

Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI

Alon Jacovi, Ana Marasović, Tim Miller, Yoav Goldberg

arXiv:2010.07487v3cs.AIcs.CY

TL;DR

The paper addresses the under-specified nature of trust in AI and formalizes Human-AI trust around vulnerability, anticipation, and explicit contracts. It distinguishes warranted from unwarranted trust, identifies intrinsic reasoning and extrinsic behavior as causes, and connects these concepts to XAI and evaluation.

  • Problem

    Trust references in XAI are vague, motivating a formal account of its prerequisites, goals, causes, and evaluation.

  • Method

    The paper builds a formalization of Human-AI trust using vulnerability, anticipation, contractual trust, trustworthiness, and intrinsic and extrinsic trust.

  • Results

    The framework connects XAI with evaluating whether models are trustworthy and whether users’ trust is warranted, while distinguishing warranted from unwarranted trust.

  • Takeaways & Limitations

    Trustworthy AI should pursue warranted trust, diagnose and avoid unwarranted trust, and assess relevant contracts and trustworthiness.

  • Takeaways & Limitations

    The paper acknowledges that warranted trust is treated as binary despite trust being dynamic and potentially multidimensional.

Abstract

from arXiv · show

Trust is a central component of the interaction between people and AI, in that 'incorrect' levels of trust may cause misuse, abuse or disuse of the technology. But what, precisely, is the nature of trust in AI? What are the prerequisites and goals of the cognitive mechanism of trust, and how can we promote them, or assess whether they are being satisfied in a given interaction? This work aims to answer these questions. We discuss a model of trust inspired by, but not identical to, sociology's interpersonal trust (i.e., trust between people). This model rests on two key properties of the vulnerability of the user and the ability to anticipate the impact of the AI model's decisions. We incorporate a formalization of 'contractual trust', such that trust between a user and an AI is trust that some implicit or explicit contract will hold, and a formalization of 'trustworthiness' (which detaches from the notion of trustworthiness in sociology), and with it concepts of 'warranted' and 'unwarranted' trust. We then present the possible causes of warranted trust as intrinsic reasoning and extrinsic behavior, and discuss how to design trustworthy AI, how to evaluate whether trust has manifested, and whether it is warranted. Finally, we elucidate the connection between trust and XAI using our formalization.

1 INTRODUCTION

The paper formalizes Human-AI trust around user vulnerability and the ability to anticipate AI decisions, then uses this framework to explain trust’s causes, evaluation, and relationship to XAI.

  • 1 INTRODUCTION: Human-AI trust requires user vulnerability and aims to enable anticipation of the impact of AI decisions.The user must be exposed to consequences of the AI’s actions, while trust targets anticipation under uncertainty.
  • 1 INTRODUCTION: Contractual trust specifies what the AI is trusted with, while warranted and unwarranted trust distinguish whether trust achieves its goal.Together, these concepts complete the paper’s formal definition of Human-AI trust.
  • 1 INTRODUCTION: Intrinsic trust arises from the AI’s observable reasoning process, whereas extrinsic trust arises from its external behavior.These mechanisms describe how an AI model can gain a person’s warranted trust.
  • 1 INTRODUCTION: The framework connects intrinsic and extrinsic trust to XAI and supports evaluation of interaction vulnerability, anticipation, and trust.The paper also discusses relations to interpersonal trust and human-machine trust, plus possible extensions.
  • 1 INTRODUCTION: The paper formalizes Human-AI trust as distinct from, but rooted in, sociological interpersonal trust.It aims to support principled development of AI that should and will be trusted in practice.

2 A BASIC DEFINITION OF TRUST

The paper defines Human-AI trust as trust under risk: users accept vulnerability to AI actions while attempting to anticipate their impact, with distrust representing protective rejection of that vulnerability.

  • 2 A BASIC DEFINITION OF TRUST: Interpersonal trust involves believing that another party will act in one’s best interest while accepting vulnerability to that party’s actions.Its goal is to make behavior predictable and collaboration easier.
  • 2 A BASIC DEFINITION OF TRUST: Human-AI trust adapts interpersonal trust by treating anticipation of behavior under risk as its central purpose.The paper emphasizes anticipation and vulnerability as the key properties carried into the AI setting.
  • 2 A BASIC DEFINITION OF TRUST: Risk is a prerequisite for Human-AI trust because the user must perceive an undesirable event as both possible and consequential.Trust can ideally be verified only after establishing both conditions.
  • 2 A BASIC DEFINITION OF TRUST: In credit scoring, trust requires the loan officer to recognize that an AI decision could be incorrect and lead to an undesirable default-related outcome.Applicants may face different risks, such as denial or higher interest on a deserved loan.
  • 2 A BASIC DEFINITION OF TRUST: Distrust is belief that the AI may not act in the user’s best interest, causing the user to reject vulnerability rather than merely lack trust.The paper treats distrust as trust in a negative scenario, not as simple absence of belief.
  • 2 A BASIC DEFINITION OF TRUST: Anticipation is the goal of Human-AI trust, not necessarily evidence that trust exists or is absent.A user’s ability to anticipate AI behavior does not by itself diagnose whether trust has manifested.

3 CONTRACTUAL TRUST

The paper makes trust explicitly contractual: users trust an AI to uphold specified behavior, and explanations and evaluations must therefore match the contract and context being considered.

  • 3.1 Trust in Model Correctness: A model matching a random baseline can still support instance-specific trust when explanations reveal sub-populations where its predictions are more likely correct.The example separates overall accuracy from users’ ability to anticipate correctness for particular inputs.
  • 3.1 Trust in Model Correctness: Trust in model correctness concerns access to patterns distinguishing correct from incorrect cases, not merely the model’s general performance.An explanation can increase predictability and instance-level trust without changing measured performance.
  • 3 CONTRACTUAL TRUST: Contractual trust is belief that an AI will adhere to a specific contract defining the behavior the user anticipates.The paper argues that Human-AI trust is contractual and that the relevant contract must be explicit.
  • 3.2 The General Case: Trust in a Contract: Contracts are conditioned on context because the same apparent task may yield different reliability for familiar and infrequent feature combinations.The paper notes that models may perform strongly near training data and poorly on less frequent cases.
  • 3.2 The General Case: Trust in a Contract: European trustworthy-AI requirements and standardized documentation can specify useful contracts and information for evaluating or increasing trust.Examples include data statements, datasheets, model cards, reproducibility checklists, fairness checklists, and factsheets.
  • 3.2 The General Case: Trust in a Contract: Different contracts require different evaluation methods and explanatory methods, so broad trust comprises multiple dimensions rather than one assessment.For example, neuron-level efficiency may matter for sustainability but not for universal design.
  • 3.2 The General Case: Trust in a Contract: The formalization clarifies that contracts specify anticipated behavior, while trusting the AI means believing that the relevant contracts will be upheld.This frames trust as a multidimensional transaction whose relevant dimensions should be explored before societal integration.

4 TRUSTWORTHY AI

Trustworthiness is a property of whether an AI can maintain a contract, whereas warranted trust is trust caused by that capability. Trust can therefore exist without trustworthiness, and interface-driven trust may be unwarranted.

  • Trustworthiness means that an AI model is capable of maintaining a specified contract.
  • Trust and trustworthiness are distinct: a model may be trustworthy without gaining trust, or gain trust without being trustworthy.
  • A high-quality interface can increase confidence without improving accuracy, making interface-caused trust unwarranted for a performance contract.
  • Warranted trust requires a causal relationship: theoretically manipulating the model’s contract-maintaining capability would change the user’s trust.
  • Unwarranted trust should be evaluated and minimized because trust exceeding trustworthiness can produce misuse, while the reverse can produce disuse.
  • Warranted distrust occurs when distrust is caused by the AI’s inability to maintain a relevant contract.

5 DEFINING HUMAN-AI TRUST

The paper defines Human-AI trust contractually: a human accepts vulnerability while believing an AI can maintain a contract. Trust aims at anticipating that maintenance under uncertainty and requires perceived risk.

  • The paper treats AI as automation attributed by the human with human-like intelligence or intent.
  • An AI is trustworthy to contract C if it is capable of maintaining C.
  • Human-AI trust exists when H perceives M as trustworthy to C, accepts vulnerability to M’s actions, and seeks to anticipate M maintaining C under uncertainty.
  • Anticipation is the user’s belief that the AI will work as intended, based on perceiving the model as capable of maintaining the contract.
  • Trust is warranted when it is caused by the model’s trustworthiness; otherwise, it is unwarranted.
  • Distrust is contractual when H perceives M as unable to maintain C and therefore rejects vulnerability to M’s actions.

6 CAUSES OF TRUST

The paper divides warranted trust into intrinsic and extrinsic causes. Intrinsic trust depends on comprehended, prior-aligned reasoning, while extrinsic trust depends on trustworthy evaluation of behavior and generalization.

  • 6 CAUSES OF TRUST: The paper divides causes of warranted trust into intrinsic and extrinsic types.
  • 6.1 Intrinsic Trust: Intrinsic trust increases when the model’s observable decision process matches the user’s priors about trustworthy reasoning.
  • 6.1 Intrinsic Trust: Explanation enables intrinsic trust only when users comprehend the model’s true reasoning and judge it consistent with agreeable reasoning priors.
  • 6.1 Intrinsic Trust: Model simplicity alone may not produce intrinsic trust when users lack the task knowledge or priors needed to assess trustworthy behavior.
  • 6.1 Intrinsic Trust: Interpretability must communicate reasoning at a granularity relevant to the user’s priors, not merely information that is accurate and comprehensible.
  • 6.2 Extrinsic Trust: Extrinsic trust is trust in an evaluation scheme that justifies generalization from evaluated behavior to future unseen instances.
  • 6.2 Extrinsic Trust: Extrinsic trust requires both a trustworthy model and a trustworthy evaluation scheme whose evaluated distribution matches relevant future unseen instances.
  • 6.2 Extrinsic Trust: Specialized test sets can target contracts such as robustness, fairness, privacy, or discrimination rather than only general post-deployment performance.

7 EXPLAINABILITY AND TRUST

The paper extends XAI’s trust motivation into a contract-specific account: explanations may increase trustworthiness, trust in trustworthy AI, or distrust in non-trustworthy AI, thereby supporting warranted responses.

  • XAI for Trust (extended): XAI should target a stated contract by increasing AI trustworthiness, user trust in trustworthy AI, or user distrust in non-trustworthy AI.These outcomes are intended to produce warranted trust or distrust in that contract.
  • Trustworthiness: AI is trustworthy to a contract when it can maintain that contract, while XAI can reveal reasoning signals relevant to that capability.Hiding relevant signals may make the model less trustworthy for some contracts.
  • Trust: The user’s goal in trust is anticipating behavior under risk, and XAI gives easier access to signals that support this anticipation.The relevant target is not trust in general but anticipation for a specific contract.
  • Distrust: XAI can also enable warranted distrust by helping users anticipate when desired behavior will not occur.Distrust is framed as the inverse of trust for the relevant contract.

8 EVALUATING TRUST

Evaluating trust requires distinguishing user vulnerability, anticipated behavior, trustworthiness, and whether trust is caused by that trustworthiness. The paper argues that many existing evaluations omit these conditions.

  • Vulnerability in Trust: Trust evaluation requires user vulnerability: the user must face realistic favorable and unfavorable outcomes from depending on the AI.Simply asking whether users trust a model for a trivial task evaluates neither trust nor trustworthiness.
  • Vulnerability in Trust: Figure 3 classifies ERASER datasets by human effort and seriousness of undesirable events; only two to three tasks fall beyond both trust-relevant boundaries.Evidence Inference is given as an example of a task in that region.
  • Vulnerability in Trust: Use-cases for trust studies should involve considerable human effort and vulnerability, with difficulty assessed beyond task duration and use-cases defined explicitly.The task domain need not exactly replicate the real interaction setting.
  • Vulnerability in Trustworthiness: Trustworthiness can be assessed without user vulnerability, such as when experts inspect whether a model performs expected reasoning steps.Accuracy evidence may support trustworthiness but does not itself evaluate trust because no risk is involved.
  • Warranted and Unwarranted Trust: A manipulationist protocol measures trust, changes the model’s real trustworthiness, and measures trust again; the resulting change indicates warranted trust.Suggested manipulations include handicapping, improving predictions, or replacing the model with an oracle.
  • Evaluating Anticipation: Simulatability can proxy successful anticipation, but it does not verify that trust exists or that trust is warranted.It measures whether users can simulate the AI’s outcome at the instance level.

9 DISCUSSION

The discussion situates Human-AI trust within interpersonal trust while distinguishing AI-specific contractual and causal notions. It also identifies multidimensionality, developer influence, and trustor attributes as extensions or boundaries.

  • Other Elements of Interpersonal Trust and Trust in Automation: The formalization begins with sociology’s interpersonal trust and relates it to trust in automation and other interpersonal-trust dimensions.The paper treats these related constructs as relevant to, but not identical with, Human-AI trust.
  • Warranted and Justified Trust in Social Science: Human-AI trust uses a stricter causal definition of warranted trust because AI capabilities are well-defined and AI lacks real intent.This differs from sociological definitions based on whether trust was betrayed.
  • Can Trust Become Warranted?: Contractual trust allows trust to be calibrated per contract, even though general trust may be multidimensional rather than a single state or scale.Contracts are treated as dimensions tied to the model’s capability to maintain them.
  • Trust in the AI Developer: Trust in the AI developer may influence user–AI interaction but is classified as interpersonal trust by proxy rather than Human-AI trust.The paper positions developer trust as future work for a more nuanced model.
  • Personal Attributes of the Trustor: The model treats each interaction as a clean-slate transaction and leaves personality, sociocultural background, prior trust, and trust restoration for future work.These personal attributes are identified as possible additions to the model.

10 CONCLUSION

The conclusion turns the formalization into design and evaluation guidance: identify risk and contracts, distinguish warranted from unwarranted trust, and connect explanations to trustworthiness and user calibration.

  • Takeaways: The paper calls for more accurate trust discussions, explicit distinctions between intrinsic and extrinsic causes, and verified risk in user studies.These guidelines are intended to support disciplined development of trustworthy and trusted AI.
  • Takeaways: Trust assessment should begin by verifying that the user faces unfavorable and realistic risks from the AI’s actions.Risk assessment is presented as necessary before assessing trust.
  • Takeaways: Developers should explicitly identify the contracts their models maintain and use suitable affordances to make those contracts clear.This helps avoid unwarranted trust based on contracts users assume but the developer does not intend to uphold.
  • Takeaways: Successful anticipation does not establish warranted trust because anticipation may depend on variables such as interface quality rather than trustworthiness.Simulatability is useful for assessing anticipation but dangerous to use alone.
  • Takeaways: Unwarranted trust can contribute to misuse, disuse, or abuse, so AI research should diagnose and avoid it by identifying contracts and assessing trustworthiness.The paper considers trust ethically desirable only when warranted.
  • Takeaways: Warranted distrust can make imperfect AI useful when achieving complete trustworthiness for all relevant contracts is prohibitively difficult.Distrust should follow non-trustworthiness for contracts relevant to the application.
Loading 2010.07487v3…