Source-linked AI summary
The relationship between trust in AI and trustworthy machine learning technologies
Ehsan Toreini, Mhairi Aitken, Kovila Coopamootoo, Karen Elliott, Carlos Gonzalez Zelaya, Aad van Moorsel
TL;DR
The paper addresses how machine-learning technologies affect trust in AI-based services and products. It systematically relates social-science trust concepts to trustworthy machine-learning technologies through ABI+ and HET qualities, organizing them into FEAS categories and a Chain of Trust across the AI life cycle. It concludes that FEAS technologies are closely related to technologies pursued in ethical AI and to international Principled AI frameworks.
Problem
Understanding how machine-learning technologies impact trust is necessary for building AI-based systems that users and the public can justifiably trust.
Method
The paper relates ABI+ trust frameworks and HET technology qualities to Fairness, Explainability, Auditability and Safety technologies across AI-system life-cycle stages and Principled AI frameworks.
Results
The paper establishes a connection between social-science trust and trustworthy machine-learning technologies, identifying FEAS categories and their relation to Principled AI frameworks.
Takeaways & Limitations
Trust-enhancing technologies should be considered across interrelated stages of an AI system’s life cycle, forming a Chain of Trust.
Takeaways & Limitations
The precise technology needs of Principled AI frameworks require deeper investigation because existing frameworks discuss issues at a coarser granularity than the computing literature.
Abstract
from arXiv · showhide
To build AI-based systems that users and the public can justifiably trust one needs to understand how machine learning technologies impact trust put in these services. To guide technology developments, this paper provides a systematic approach to relate social science concepts of trust with the technologies used in AI-based services and products. We conceive trust as discussed in the ABI (Ability, Benevolence, Integrity) framework and use a recently proposed mapping of ABI on qualities of technologies. We consider four categories of machine learning technologies, namely these for Fairness, Explainability, Auditability and Safety (FEAS) and discuss if and how these possess the required qualities. Trust can be impacted throughout the life cycle of AI-based systems, and we introduce the concept of Chain of Trust to discuss technological needs for trust in different stages of the life cycle. FEAS has obvious relations with known frameworks and therefore we relate FEAS to a variety of international Principled AI policy and technology frameworks that have emerged in recent years.
1 Introduction
This paper connects social-science trust frameworks with machine-learning technologies that may enhance trust in AI-based services and products. It classifies these technologies as Fairness, Explainability, Auditability and Safety (FEAS), relates them to Principled AI frameworks, and considers trust across the AI life cycle.
- Research aim: The paper examines how machine-learning technologies might enhance or impact trust, drawing on social-science and trust-in-technology frameworks.It distinguishes trust in AI from the ethics of AI: trustworthy AI concerns technological qualities, whereas trust is a response to technologies or development processes.
- Trust framework: The authors use ABI+ and HET qualities to identify technological qualities that can support trust in AI-based systems.ABI+ builds on Ability, Benevolence and Integrity, adds predictability, and incorporates a temporal dimension from initial to continuous trust.
- AI life cycle: The paper introduces an AI Chain of Trust linking trust enhancement across stages of an AI system’s life cycle.It addresses design, development and deployment phases and the interrelations among life-cycle stages.
- FEAS classification: FEAS classifies trust-enhancing machine-learning technologies into Fairness, Explainability, Auditability and Safety.The classification is presented as the paper’s central technology-oriented contribution.
- Principled AI frameworks: The FEAS classification is related to Principled AI frameworks used in ethics and policy discussions.The paper discusses this relationship alongside a wider set of technologies than those typically derived from implementing Principled AI frameworks.
2 Trust
The paper frames trust as a context-dependent, evolving relationship rather than a binary judgment, and distinguishes trust in AI from ethical evaluations of AI technologies. It connects social-science trust frameworks with technological qualities that can shape perceptions of trustworthiness.
- 2.1 The ABI Framework: Ability, Benevolence and Integrity: The ABI framework identifies Ability, Benevolence, and Integrity as the main attributes shaping assessments of a party’s trustworthiness.Ability concerns domain-specific skills and competencies; Benevolence concerns wanting to do good for the trustor; Integrity concerns adherence to acceptable principles.
- 2.1 The ABI Framework: Ability, Benevolence and Integrity: The ABI+ model adds Predictability or Reliability, emphasizing that trustworthiness judgments are sustained through ongoing relationships.Predictability reinforces perceptions of Ability, Benevolence, and Integrity over time.
- 2.1 The ABI Framework: Ability, Benevolence and Integrity: Trust is conditional and exists along a continuum rather than as a simple trust-or-distrust state.Assessments depend on the task, circumstances, contextual factors, and the trustor’s predispositions.
- 2.2 Trust in Science and Technology: AI-based services combine scientific and technological characteristics, making public trust in science and technology relevant to understanding trust in AI.The paper notes that reliance on expert knowledge and behavior makes trust in AI increasingly conditional.
- 2.3 Trust Technology Qualities: Humane, Environmental and Technological: Siau et al. identify humane, environmental, and technological qualities as conditions through which technologies may be perceived as trustworthy.Humane qualities depend partly on personality, past experience, and cultural background, including whether testing a product or service is feasible.
- 2.4 Time-domain: Initial and Continuous Trust: Trust develops through initial and continuous phases, with events such as data breaches, privacy leaks, and unethical-practice reports affecting later trust levels.Continuous trust reflects changes in the ongoing trustor–trustee relationship after initial conditions are established.
3 Machine Learning and the Chain of Trust
The paper presents machine learning as a staged pipeline and introduces the Chain of Trust to connect trust considerations across those stages and over time. Trust effects can propagate between stages, emerge later, recur through iteration, and arise across design, development, deployment, and failure events.
- 3.3 Chain of Trust: The Chain of Trust connects trust considerations across stages of the machine learning pipeline and can expand into a cycle as services iterate over time.The concept links pipeline stages with the lifecycle of an AI-based service or product.
- 3.1 Basics of Machine Learning: A machine learning algorithm maps feature inputs x to outputs y through a parameterized function y = f_θ(x).The parameters θ are tuned using a predefined loss function for the model’s task.
- 3.1 Basics of Machine Learning: Supervised, unsupervised, and reinforcement learning differ in how inputs and outputs are labeled or selected for the task.Supervised learning assigns predefined labels, unsupervised learning detects hidden patterns, and reinforcement learning selects actions, observations, or rewards.
- 3.2 Machine Learning Pipeline: The pipeline includes data collection, data preparation, feature extraction, training, testing, and deployment stages.Data-centric stages prepare inputs, while model-centric stages tune, evaluate, and deploy the model.
- 3.3 Chain of Trust: Improved data preparation may reduce bias in algorithmic outputs, but its trust impact may become visible only through later results or communicated evidence.The paper gives better outcomes, news reports, and long-term statistics as ways such effects may become visible to users.
- 3.3 Chain of Trust: Trust effects can propagate between stages, recur when new data triggers retraining and inference, and arise throughout design, development, and deployment.The paper also identifies accidents such as security breaches, data loss, and discovered bias as possible trust failures across pipeline and lifecycle stages.
4 Trust in AI-Based Systems
The section connects trust in AI-based solutions with trustworthy underlying technologies, distinguishing model and data concerns and proposing FEAS technologies as a technology-oriented classification. It also relates these technologies to Principled AI frameworks while noting that policy frameworks often lack technical specificity.
- Trust in AI-based solutions depends on connecting social concerns with the trustworthiness of underlying machine learning technologies.The section frames this connection as the basis for examining trust-related issues in AI services and products.
- AI Chain of Trust: The paper introduces an AI Chain of Trust to address trust-enhancing technologies across interrelated stages of an AI system life cycle.The section distinguishes technologies that verify model outcomes from approaches that redesign models or algorithms to be inherently more trustworthy.
- Data-related trust concerns: Data-related concerns span collection, preprocessing, storage, privacy, consent, control, and potential illegitimate use of personal data.GDPR illustrates requirements such as transparency, explicit consent, control over collected data, and the right to be forgotten.
- Model-related trust concerns: Model-related concerns include unfairness, non-transparency, lack of accountability, and the limits of optimizing primarily for accuracy and efficiency.Accuracy and speed may demonstrate ability, but do not necessarily satisfy other trust qualities; algorithms are often treated as black boxes.
- Principled AI frameworks: Principled AI frameworks broadly relate to FEAS technologies, but their high-level qualities do not specify technical needs at the granularity of computing research.For fairness alone, the technology literature contains at least 21 mathematical definitions and many bias-prevention or detection approaches.
- FEAS technologies: FEAS classifies trustworthy machine learning technologies as Fair, Explainable, Auditable, and Safe, alongside essential algorithmic accuracy and efficiency.The classification is intended to align technology discussions with existing Principled AI frameworks.
5 Trustworthy Machine Learning Technologies
The paper groups trustworthy machine learning approaches into Fair, Explainable, Auditable, and Safe technologies, describing representative methods and remaining challenges. These technologies address discrimination, model understanding, provenance, security, privacy, and adversarial threats.
- 5.1 Fairness Technologies: Fairness technologies seek non-discriminating outcomes, but fairness is difficult to measure because the literature contains at least 21 definitions that cannot necessarily all be achieved simultaneously.Trust-enhancing fairness metrics must also relate technical measurements to impacts on individuals and the public.
- 5.1 Fairness Technologies: Fairness approaches include dataset debiasing, decision-value adjustments, and model-specific regularization, while removing protected features alone is insufficient or potentially unhelpful.The paper presents these as examples rather than a complete review of fairness techniques.
- 5.2 Explainability Technologies: Explainability relates model operations and outcomes to human-understandable terms through ex-ante intelligible models or ex-post explanatory models.Ex-post explainers are categorized as local or global.
- 5.2 Explainability Technologies: Explainability remains constrained by trade-offs between fidelity to the original model and the intelligibility of explanations.The paper identifies this trade-off as an open challenge for explainable machine learning technologies.
- 5.3 Auditability Technologies: Auditability enables third parties to challenge model operations and outcomes and uses lineage or decision provenance to expose data and system history.Decision provenance can provide the history of particular data and a view of system behavior and interactions with internal or external entities.
- 5.4 Safety Technologies: Machine learning data and models can face targeted, indiscriminate, stealthy, or exploratory adversarial attacks.Security and privacy protections draw on confidentiality and integrity, including differential privacy and homomorphic encryption.
6 Conclusion
The conclusion links social-science trust concepts with trustworthy machine learning technologies through FEAS categories and an AI Chain of Trust. It also reports a close relationship between FEAS and international Principled AI frameworks.
- The paper connects the ABI framework and HET technology qualities with Fair, Explainable, Auditable, and Safe machine learning technologies.
- FEAS technologies should be considered across interrelated stages of an AI system life cycle, with each stage contributing to a Chain of Trust.
- The paper maps FEAS technologies onto concerns in a large set of international Principled AI policy and technology frameworks.