Source-linked AI summary
A Comprehensive Taxonomy for Explainable Artificial Intelligence: A Systematic Survey of Surveys on Methods and Concepts
Gesina Schwalbe, Bettina Finzel
TL;DR
XAI research has accumulated many methods, terminologies, and evaluation criteria, creating a need for a unified taxonomy that supports comparison and use-case-oriented selection. The paper conducts a structured survey-of-surveys, merges concepts into a taxonomy, and illustrates it with diverse categorized methods. It provides a broad reference for beginners, researchers, and practitioners while acknowledging that its literature search and broadly applicable structure cannot remain complete for every XAI sub-field.
Problem
The growing variety of XAI methods and overlapping taxonomies makes it difficult to grasp the field, compare methods, and select methods based on use-case requirements.
Method
The paper systematically reviews surveys on XAI methods, metrics, and method traits, merges their terminologies and concepts, and illustrates the resulting taxonomy with categorized example methods.
Results
The paper unifies scattered literature into an overarching taxonomy organized around the task, the explainer, and evaluation metrics, with applicability evidenced through method examples.
Takeaways & Limitations
The taxonomy serves as a wide-ranging reference and starting point for beginners, researchers, and practitioners working with XAI method traits and aspects.
Takeaways & Limitations
The survey is not complete because its structured literature search is biased by current sub-field interests and search terms, and sub-fields may require different or more detailed taxonomies.
Abstract
from arXiv · showhide
In the meantime, a wide variety of terminologies, motivations, approaches, and evaluation criteria have been developed within the research field of explainable artificial intelligence (XAI). With the amount of XAI methods vastly growing, a taxonomy of methods is needed by researchers as well as practitioners: To grasp the breadth of the topic, compare methods, and to select the right XAI method based on traits required by a specific use-case context. Many taxonomies for XAI methods of varying level of detail and depth can be found in the literature. While they often have a different focus, they also exhibit many points of overlap. This paper unifies these efforts and provides a complete taxonomy of XAI methods with respect to notions present in the current state of research. In a structured literature analysis and meta-study, we identified and reviewed more than 50 of the most cited and current surveys on XAI methods, metrics, and method traits. After summarizing them in a survey of surveys, we merge terminologies and concepts of the articles into a unified structured taxonomy. Single concepts therein are illustrated by more than 50 diverse selected example methods in total, which we categorize accordingly. The taxonomy may serve both beginners, researchers, and practitioners as a reference and wide-ranging overview of XAI method traits and aspects. Hence, it provides foundations for targeted, use-case-oriented, and context-sensitive future research.
1 Introduction
The paper addresses the difficulty of understanding black-box machine-learning models and selecting suitable XAI methods across diverse use cases. It responds with a structured taxonomy derived from a broad survey of surveys and illustrated with categorized method examples.
- Black-box machine-learning models can achieve high performance while hiding their learning process, internal representation, and final processing from humans.
- XAI methods are motivated by legal, safety, fairness, security, and other public-interest or business requirements.
- Choosing an appropriate XAI method requires use-case analysis and knowledge of method traits such as portability and locality.
- The paper aims to support beginners, practitioners, and researchers through a complete taxonomy of XAI method aspects with illustrative method examples.
- Its contributions include a detailed taxonomy, a survey-of-surveys covering more than 50 works, and a review of more than 50 diverse XAI methods.
- The paper reviews related work, presents its systematic review, develops the taxonomy, illustrates aspects with example methods, and concludes the study.
2 Background
The background situates XAI within its historical development, related surveys, and foundational terminology. It distinguishes the paper’s method-focused taxonomy from use-case analysis and introduces concepts such as understanding, explicability, explainability, and transparency.
- 2.1 Related work: Existing XAI meta-studies are often brief, whereas this work examines a broad collection of surveys and considers their focus, detail, and usefulness for different audiences.
- 2.1 Related work: Prior XAI taxonomies are frequently shallow, sub-field-specific, or used mainly as survey chapter structures; existing method-trait discussions also contain unique, non-overlapping aspects.
- 2.1 Related work: Use-case and requirements analysis is outside the paper’s scope; instead, it focuses on XAI method aspects that can inform requirements, such as model agnosticism and information amount.
- 2.2 History of XAI: XAI developed from earlier ideas about declarative and transparent AI toward modern efforts to make neural-network decisions understandable to stakeholders.
- 2.2 History of XAI: XAI seeks explainable machine-learning models with high learning performance and a user-centric approach that helps humans understand artificial systems.
- 2.2 History of XAI: The field increasingly treats explanation quality as something that must be evaluated using formalized measures, not merely generated.
- 2.3 Basic definitions: Understanding includes mechanistic and functional forms, while explicability makes model properties inspectable and explainability makes reasoning, models, or decision evidence accessible to humans.
- 2.3 Basic definitions: Global explanations address a model or its logic as a whole, whereas local explanations address individual model decisions or predictions.
3 Approach to literature search
The paper conducts a broad, systematic search for XAI surveys and taxonomies, then categorizes selected works by focus, length, audience, reception, and recency. It reviews over 70 surveys and uses more than 50 selected works and example methods to structure the resulting overview.
- The literature analysis targeted papers from 2010 to 2021 covering XAI method reviews, metrics, and taxonomy aspects.
- The search used two iterations: one for XAI taxonomies and one for general XAI surveys that might implicitly use taxonomies.
- The taxonomy search combined machine-learning, explainability, and taxonomy-related terms, scanning the first 300 Google Scholar results by title before abstract screening.
- Additional categorization criteria included length, target audience, citation count per year, and recency, with citation databases used as reception proxies.
- General focus: Reviews were categorized by general focus into general method collections, domain-specific collections, conceptual reviews, and toolboxes.
- Results: The review included over 70 surveys, while a sub-selection of more than 50 cited or generally interesting surveys and more than 50 diverse XAI methods received detailed treatment.
4 A survey of surveys on XAI methods and aspects
The paper surveys a broad and growing body of XAI surveys, organizing them by focus, audience, detail, and application. This synthesis supports finding relevant literature and informs a more complete taxonomy of XAI methods.
- The review addresses the difficulty of navigating abundant XAI surveys with differing focuses, detail levels, and target audiences.These needs are especially relevant for beginners and practitioners seeking a suitable entry point.
- The authors analyze more than 50 selected surveys for their key focus points, level of detail, and intended audience.The search emphasizes works proposing or using structured views of XAI methods and metrics.
- The meta-study reports that its analysis of survey taxonomies leads to what the authors describe as the most complete XAI method taxonomy available so far.This result connects the survey-of-surveys to the taxonomy developed later in the paper.
- Reviewed surveys are grouped into broadly focused works, stakeholder and HCI perspectives, metric-focused surveys, and restricted-focus collections.The paper also clusters surveys by application domain, task or input type, surrogate model, and other method traits.
- Broad method collections: Broad method collections are ordered first by length as a proxy for information amount and then by citations per year as a reception measure.The paper notes that citations per year strongly correlate with survey age.
- Survey landscape: The surveyed literature includes beginner-friendly introductions, extensive holistic reviews, specialized counterfactual surveys, and domain-focused collections.Examples cover supervised learning, single-prediction explanations, rule-based systems, Bayesian networks, and counterfactual reasoning.
5 Taxonomy
The taxonomy unifies terminology and concepts from more than 70 reviewed surveys into a structured account of XAI methods and evaluation aspects. It organizes the field procedurally around problem definition, explanator properties, and metrics, while illustrating aspects with example methods.
- The taxonomy unifies terms and notions from more than 70 surveys and systematically structures them according to practical considerations.The resulting organization is intended to differentiate and evaluate XAI methods and provide a complete state-of-the-art overview.
- At the root level, the taxonomy follows the procedural steps for building an explanation system.This establishes the structure for the chapter and the taxonomy’s first two levels.
- Problem definition: Problem definition covers task traits and the interpretability of the explanandum.It is presented as the first step in designing an explanation system.
- Explanator properties: Explanator properties are divided into input, output, user interactivity, and further formal constraints.These properties describe functional and architectural characteristics of the explanator.
- Metrics: Metrics evaluate explanation-system qualities and are classified as functionally grounded, human-grounded, or application-grounded.The distinction follows dependence on subjective human evaluation and the application context.
- Requirements derivation: Requirements should be derived from taxonomy aspects and aligned with the task, explanandum interpretability needs, explanator constraints, and metric targets.The motivating goals may include verifying fairness, safety, or security; discovering knowledge; or promoting user adoption and trust.
5.1 Problem definition
Defining an XAI problem requires specifying the task and the explanandum model, while recognizing that methods are constrained by task, input, and sometimes architecture. Task-specific methods can sometimes be extended across related prediction settings by treating added dimensions as snippets.
- Problem definition: An XAI problem must specify the task being explained and the interpretability level of the explanandum model.The explanandum is the model used to solve the original task.
- Problem definition: Out-of-the-box XAI methods usually support only particular task types and input data types.White-box methods may impose additional architectural constraints.
- Task: Typical task categories include clustering, regression, classification, detection, and semantic or instance segmentation.Examples of regression predictions include bounding-box dimensions in object detection.
- Task: Classification-oriented XAI methods can often extend to detection, segmentation, and temporal settings by expressing the question over spatial or temporal snippets.Some classifier methods require continuous prediction scores rather than final discrete labels.
- Examples: RISE explains image classification through perturbation-based heatmaps, while D-RISE extends the approach to detection using distances between prediction vectors.Both methods measure how dimming input regions influences model predictions.
- Examples: LIME applies to image and text inputs by locally approximating the explanandum with a linear model over input feature snippets.The method is model-agnostic and uses snippets as the explanatory units.
5.1.2 Model interpretability
Model interpretability concerns whether the explanandum supports human understanding, and can be achieved intrinsically, through blended design, self-explanations, or post-hoc helper models. The taxonomy distinguishes transparent model forms and explanation-generating outputs.
- Model interpretability: Model interpretability means that the explanandum supports both mechanistic transparency and human functional understanding.The explanandum is the model used to solve the system’s original task.
- Interpretability strategies: Interpretability can be built intrinsically, blended into a partly interpretable model, added through self-explanations, or obtained with a post-hoc helper model.These choices differ in whether the trained explanandum model itself is changed or designed for interpretability.
- Transparency levels: Transparent models may be simulatable as wholes or decomposable into parts that humans can understand separately.Simulatability can be assessed by model size or required computation length.
- Interpretable models: Inherently interpretable model families include decision trees, rules, Bayesian models, linear or logistic models, support vector machines, and additive models.General linear and additive models can provide feature-importance weights under their respective structural assumptions.
- Interpretable models: General linear models assume transformed inputs relate linearly to expected outputs, whereas general additive models sum transformed feature contributions.These structural assumptions yield interpretable feature weights.
- Blended models: Blended models integrate transparent symbolic components with non-transparent subsymbolic components, supporting hybrid neuro-symbolic designs.Logic Tensor Networks use fuzzy logic constraints with DNN predicates and simple linear representations of relations.
- Self-explaining models: Self-explaining models produce additional outputs such as attention maps, disentangled representations, or textual and multimodal explanations.Attention maps highlight relevant input regions, while disentangled representations encode symbolic concepts in intermediate dimensions.
- Post-hoc explanations: Proxy and surrogate models are helper models that may fully mimic the explanandum or approximate selected sub-aspects such as input attribution.Training a full approximation is also called model distillation, student-teacher learning, or model induction.
5.2 Explanator
The explanator is treated as an explanation-producing function. Its taxonomy is organized around what it receives, produces, mathematically permits, and how it processes explanations.
- Explanator: An explanator can be modeled as an implemented function that outputs explanations.This abstraction supports a systematic description of explanator characteristics.
- Explanator: The explanator’s defining aspects are its input, output, function class, and explanation-processing behavior.Function class includes mathematical properties or constraints, while processing can include interactivity.
- Explanator: Figure 5 summarizes the explanator aspects discussed in the section.
5.2.1 Input
Explanator input characteristics include required resources, portability across model types, and whether explanations are local or global. The taxonomy also distinguishes model-access requirements and validity ranges.
- Input requirements: Explanators may require the model, valid data samples, user feedback, or additional situational context.The required inputs differ across methods.
- Portability: Portability measures how dependent an explanation method is on access to the explanandum model’s internals.This characteristic is also called translucency or transferability.
- Portability: Model-agnostic methods require only model inputs and outputs, whereas model-specific methods require internal processing or architectural access.LIME and SHAP exemplify model-agnostic methods; gradient-based attention methods are model-specific.
- Portability: Hybrid methods occupy an intermediate position between model-agnostic and model-specific explanation.DeepRED extracts rules through layer outputs of a DNN and therefore is neither fully model-agnostic nor fully dependent on internals.
- Locality: A surrogate’s validity range determines whether its input must represent local behavior around samples or global behavior across the input space.Local explanations address why particular decisions were made; global explanations address how decisions are made.
- Locality: SpRAy derives global behavioral patterns by spectrally clustering local LRP attribution heatmaps.This connects local feature attributions to analysis of broader model behavior.
5.2.2 Output
The taxonomy characterizes XAI output by what is explained, how it is explained, and how it is presented. It covers explanation objects ranging from training and data to processing and inner representations, alongside output forms such as examples, counterfactuals, prototypes, and feature importance.
- XAI output is characterized by the explanation object, output type, and presentation form.
- Explanation objects include development during training, data, model processing, and inner representations.
- Processing: Processing-focused methods explain decision boundaries and feature attribution using approaches such as RISE, LIME, LRP, decision trees, and rules.
- Inner representation: Inner-representation methods analyze latent spaces, layers, units, or vectors, including Feature Visualization, NetDissect, Concept Completeness, and IIN.
- Development: Training-focused methods inspect model evolution and the effects of new samples, with examples including depth analysis and Influence Functions.
- Data: Data-focused methods support visualization, feature mining, and prototype- or example-based explanations through approaches such as PCA, t-SNE, and clustering.
- Output types include feature-importance maps, rules, example instances, contrastive or counterfactual examples, and prototypes.
- Contrastive / counterfactual / near miss examples: Counterfactual explanations show how input features must change to obtain an alternative output, while CEM also identifies features that must minimally be absent.
5.2.3 Interactivity
XAI interaction can be static or interactive. Interactive systems let users inspect explanations, provide corrective feedback, and iteratively query information or combine methods through dialogue.
- Interaction is static when an explanation is presented once and interactive when the process accepts user feedback as explanation input.
- Interaction task: Users can inspect explanations by exploring parts of one explanation or considering alternatives and complementary explanations.
- Interaction task: Users can correct explanations by changing labels, re-weighting or adapting features, or imposing constraints on generated verbal explanations.
- Explanation process: Sequential analysis supports iterative queries that help users understand model decisions over time in line with their capabilities and context.
- Explanation process: Multi-modal explanations combine methods and involve users through dialogue, including phrase-critic models and explanatory dialogues.
5.2.4 Mathematical constraints
Mathematical and architectural constraints formalize desirable properties of explanators, including simplicity, satisfiability, limited query requirements, and reduced complexity.
- Mathematical constraints encode formal explanator properties considered helpful for receiving explanations.
- Linearity and monotonicity are presented as desirable forms of simplicity for proxy-model outputs.
- Satisfiability means that explanator outputs readily support applying formal methods such as solvers.
- A limited number of iterations can be desirable because repeated queries may be costly or restricted in some use cases.
- Explanator types and surrogate models can impose architectural constraints such as size or disentanglement that correlate with reduced complexity.
- Sparsity can reduce complexity in linear models, while tree depth and width can be constrained; regularization is one way to obtain sparsity.
5.3 Metrics
The paper organizes XAI evaluation metrics by required human involvement into functionally-grounded, human-grounded, and application-grounded categories. These metrics assess formal behavior, human understanding, and application-relevant performance.
- XAI metrics are categorized as functionally-grounded, human-grounded, or application-grounded according to human involvement.
- Functionally-grounded metrics: Functionally-grounded metrics measure formal explanator properties without human feedback, including complexity, stability, consistency, sensitivity, expressiveness, faithfulness, localization accuracy, completeness, overlap, and surrogate accuracy.
- Functionally-grounded metrics: Faithfulness measures how accurately an explanator conforms to the object of explanation, while greater simplification can involve a fidelity-interpretability trade-off.
- Functionally-grounded metrics: Localization accuracy measures how well explanations identify ground-truth points of interest, such as certainty, bias, feature importance, or outliers.
- Functionally-grounded metrics: Completeness measures the input-space range where high fidelity is expected, whereas overlap measures samples satisfying more than one rule.
- Functionally-grounded metrics: Architectural complexity uses properties such as feature counts, changed-feature counts, model sparsity, width, and depth to approximate perceived complexity.
- Human-grounded metrics: Human-grounded metrics use human feedback or observed reactions, often through proxy tasks rather than expensive experts or application runtime.
- Human-grounded metrics: Interpretability concerns how well a mental model approximates the explanator, while effectiveness concerns how well a person can simulate the object of explanation.
6 Discussion and conclusion
The paper unifies XAI method aspects into a taxonomy and survey of surveys intended to help diverse audiences navigate the expanding field and select methods according to use-case needs. The authors note that the survey is necessarily incomplete and broadly scoped, while future updates and practical selection guides remain desirable.
- Discussion and conclusion: The survey of surveys found expanding breadth in XAI application fields and method types, alongside exponentially increasing numbers of methods and surveys.Persistent trends include medicine, recommendation systems, and visual explanations; emerging or re-awakening trends include natural language processing and rule-based methods.
- Discussion and conclusion: The unified taxonomy organizes XAI explanations around the task, explainer, and evaluation metrics, with numerous example methods categorized by detailed aspects.The examples are also summarized using seven prominent criteria, including interpretability form, model dependence, global versus local scope, explanation object, presentation form, and explanation type.
- Discussion and conclusion: XAI explanation systems should be tailored to the use-case across development, application, and evaluation, considering stakeholder contexts and needs.The taxonomy supports analysis of stakeholder needs and formulation of use-case-specific requirements, while the survey helps identify methods that satisfy them.
- Discussion and conclusion: The taxonomy and survey provide beginners with an entry point, practitioners with method-selection support, and researchers with tools for positioning work and identifying gaps.These intended uses span overview, requirement formulation, suitable-method discovery, and research-gap analysis.
- Discussion and conclusion: The survey is not complete because its structured literature search is biased by current interests in AI sub-fields and by the selected search terms.The authors nevertheless state that this bias may preserve relevance for a large part of the research community.
- Discussion and conclusion: The taxonomy is broadly applicable, whereas individual XAI sub-fields may require more detailed aspects or a different structure.The authors anticipate future updates of the research state and practical guides for choosing methods from use-case definitions.
Declarations
The declarations report the study’s funding and employment context, doctoral-research supervision, and absence of competing financial interests.
- Declarations: Funding was obtained as declared in the acknowledgments, with the first author employed at Continental Automotive GmbH and the second at the University of Bamberg.The work was authored within both authors’ doctoral research, supervised by Prof. Dr. Ute Schmid.