Source-linked AI summary
The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations
A. Chatzimparmpas, R. Martins, I. Jusufi, K. Kucher, Fabrice Rossi, A. Kerren
TL;DR
Complex and critical ML applications can be difficult to understand and trust, motivating a systematic survey of visualization-based approaches. The paper analyzes 200 publications, builds a fine-grained trust categorization, and identifies patterns and research opportunities, while noting limits in evaluation and coverage of ML types.
Problem
ML models’ black-box nature makes their results difficult to understand and trust in complex applications, creating a need to map visualization approaches for improving trustworthiness.
Method
The STAR synthesizes 200 papers into a fine-grained categorization of trust across interactive ML facets and analyzes temporal, topic, correlation, and dataset patterns.
Results
The survey comprises 8 high-level aspects, 18 category groups, and 119 categories, alongside topic and correlation analyses of the literature.
Takeaways & Limitations
The mapping and interactive browser support researchers and practitioners in exploring visualization techniques and open directions for enhancing trust in ML models.
Takeaways & Limitations
The literature underrepresents ML types beyond classification and regression, while visualization evaluation remains difficult and beginner-focused tools are limited.
Abstract
from arXiv · showhide
Machine learning (ML) models are nowadays used in complex applications in various domains, such as medicine, bioinformatics, and other sciences. Due to their black box nature, however, it may sometimes be hard to understand and trust the results they provide. This has increased the demand for reliable visualization tools related to enhancing trust in ML models, which has become a prominent topic of research in the visualization community over the past decades. To provide an overview and present the frontiers of current research on the topic, we present a State-of-the-Art Report (STAR) on enhancing trust in ML models with the use of interactive visualization. We define and describe the background of the topic, introduce a categorization for visualization techniques that aim to accomplish this goal, and discuss insights and opportunities for future research directions. Among our contributions is a categorization of trust against different facets of interactive ML, expanded and improved from previous research. Our results are investigated from different analytical perspectives: (a) providing a statistical overview, (b) summarizing key findings, (c) performing topic analyses, and (d) exploring the data sets used in the individual papers, all with the support of an interactive web-based survey browser. We intend this survey to be beneficial for visualization researchers whose interests involve making ML models more trustworthy, as well as researchers and practitioners from other disciplines in their search for effective visualization techniques suitable for solving their tasks with confidence and conveying meaning to their data.
1. Introduction
Trust in ML models is difficult to establish in complex, high-stakes applications, motivating visualization research and this STAR’s systematic mapping of the field.
- Motivation: ML models increasingly support complex and critical decisions, making their trustworthiness especially important in domains such as medicine and criminal justice.The paper highlights decisions involving human lives as a key context for trustworthy models.
- Motivation: Explainable AI initiatives identify explainability, appropriate trust, and effective management of AI as important future-development goals.The paper presents DARPA’s XAI program as one response to challenges surrounding trust.
- Related work: Visualization research has produced tools and techniques for interpreting ML models, supporting collaboration, and making analysts more comfortable with ML solutions.Examples include visual exploration of predictive models and global or local interpretability approaches.
- Contributions: The STAR maps visualization techniques, reported effectiveness, application domains, meanings of trust, and open research challenges in this literature.Its contributions include an empirically informed trust definition and a fine-grained categorization derived from 200 papers.
- Contributions: An interactive online survey browser supports the report’s categorization, pattern identification, and reader-led investigation of the literature.The browser is presented as a tool for exploring the report’s data and findings.
2. Background: Levels of Trustworthiness of Machine Learning Models
The paper frames trustworthiness as a multi-level property of the ML pipeline and develops trust categories informed by prior work and practitioner input, while reporting varied expectations for visualization.
- Practitioner input: A questionnaire of 27 ML practitioners found broadly positive attitudes toward visualizing data sources, quality, algorithm comparisons, tuning, what-if scenarios, and fairness.Participants generally selected scores of 4 or 5, while rejecting the sufficiency of a single performance metric without further visualization.
- Practitioner input: Participants’ open-ended suggestions favored visualizing the ML process across phases and highlighted feature importance among specific concepts.Feature importance received four mentions in the reported responses.
- Levels of trustworthiness: Trustworthiness is divided into five levels corresponding to stages of a typical ML pipeline, from raw-data collection through later model-related processes.The levels reflect both increasing abstraction and the sequential structure of the pipeline.
- Levels of trustworthiness: Trust issues relevant to multiple levels are assigned to the lowest applicable level because weaknesses can cumulatively destabilize model predictions.The categorization treats trust concerns as cascading across pipeline stages.
- Trust categories: The raw-data level includes source reliability, transparent collection, and data bias, including subgroup-specific differences that may produce unfair decisions.Visualization can expose collection sources, reliability indicators, and potential subgroup bias.
- Trust categories: The learning-method level covers familiarity, interpretability, explainability, knowledgeability, and fairness in relation to ML algorithms.Visualization may compare models, explain algorithm details, and support users who lack algorithm familiarity or visualization literacy.
3. Related Surveys
Existing surveys address interpretability, explanation, debugging, predictive visual analytics, deep learning, clustering, or dimensionality reduction, but none explicitly centers on categorizing and analyzing visualization techniques for trust in ML models. This STAR adopts a broader trust-focused perspective across the ML pipeline and user interaction.
- Existing survey literature did not explicitly categorize and analyze visualization techniques centered on trust in ML models.Related surveys discussed accuracy, quality, errors, stress, and uncertainty, but trust was not their explicit organizing focus.
- Earlier survey work covered ML visualization tools through understanding, diagnosis, and refinement, whereas this STAR focuses specifically on enhancing trust.
- Interactive Machine Learning: Because steering and refining models with domain knowledge can introduce cumulative biases, the STAR explicitly analyzes user-related biases.
- Predictive Visual Analytics: Predictive visual analytics surveys describe data, feature, training, selection, validation, visualization, and adjustment stages but do not analyze trust issues added across the pipeline.
- Deep Learning: The STAR extends beyond deep-learning interpretability by considering model bias–variance trade-offs and in situ structural comparisons across ML models.
- Deep Learning: Its categorization incorporates visualization for ML processing because prior deep-learning surveys reported that few tools visualize training processes rather than only ML results.
- Clustering and Dimensionality Reduction: The survey draws on prior clustering and dimensionality-reduction taxonomies but retains only the linear-versus-nonlinear distinction for its own categorization.
4. Methodology of the Literature Search
The STAR used a broad, manually executed literature search across visualization and machine-learning venues, followed by keyword validation, snowballing, screening, and reviewer adjudication. The process produced a 200-paper collection while explicitly excluding work that did not focus on trust in ML models.
- Search Strategy: The search combined trust- and visualization-related keywords with machine-learning terms into pairs for venue searches.A validation process scanned for additional papers and handled questionable cases.
- Search Strategy: The authors manually searched publications from January 2008 through January 2020, beginning with visualization venues and extending to established ML venues.
- Search Scope: The search covered visualization journals, conferences, workshops, online libraries, and ML venues including ICML, KDD, and ESANN.
- Scope Boundaries: The search used broad keyword combinations that produced irrelevant results requiring removal in a subsequent methodology phase.IEEE TVCG and IEEE VAST together yielded around 750 publications.
- Validation and Screening: Snowballing related-work sections added relevant papers from venues including Neurocomputing, IEEE Transactions on Big Data, ACM TIST, ECCV, and CVM.
- Validation and Screening: Papers were screened through titles, abstracts, visualizations, two-reviewer assessment of uncertain cases, and third-reviewer decisions when reviewers disagreed.
- Validation and Screening: Less than 20% disagreement occurred among 70 uncertain cases, and the final screening process retained 200 papers.
- Scope Boundaries: The survey excluded visualization techniques that did not explicitly support trust in ML models, including work focused exclusively on input-data labeling.
5. General Overview of the Relations Between the Papers
The paper analyzes temporal growth, publication venues, and co-authorship among 200 collected papers to characterize the research landscape and identify collaboration patterns. Interest increased sharply in 2018 and 2019, while the venue distribution suggests limited cross-community collaboration.
- The overview combines temporal analysis, venue analysis, and co-authorship-network analysis to examine the collected literature and its relationships.
- Time and Venues: The collection contains 200 papers, with stable interest growth since 2009 and a sharp increase in publications during 2018 and 2019.The report also describes promising numbers for 2020, whose data collection ended in January.
- Time and Venues: Many workshops co-located with ML venues indicate efforts to reach across visualization and machine-learning communities.
- Time and Venues: The small number of publications outside visualization venues may indicate difficulty for visualization researchers in collaborating with ML experts.
- Time and Venues: Table 2 reports technique counts by publication venue, separating visualization venues from other disciplines and marking journals with J and workshops with W.
- Co-authorship Analysis: The co-authorship analysis uses node size to represent authors’ in-degree values and colors the top eight clusters.
- Co-authorship Analysis: The co-authorship network is intended to reveal missing connections and potential collaboration gaps between visualization and ML communities.
6. In-Depth Categorization of Trust Against Facets of Interactive Machine Learning
The STAR develops a multifaceted categorization of trust against interactive ML facets, using 200 papers and an interactive survey browser to expose relationships and research opportunities.
- Designing and filling the categorization: Its inputs combine prior surveys, iterative paper selection, and feedback from an online questionnaire.
- The categorization contains 8 overarching aspects, 18 category groups, and 119 individual categories.
- Categorization structure: The hierarchy covers data, machine learning, processing phase, treatment method, visualization, evaluation, trust levels, and target group.
- Novel categories: The authors add categories for model-agnostic versus model-specific methods, visual granularity, verbalization, visualization evaluation, and trust levels.
- The categorization aims to reveal relationships between trust and other categories through later statistical and correlation analyses.
6.1. Data
The Data aspect links application inputs and target variables to trust-enhancing visualization techniques, while the surveyed literature spans varied domains and model types.
- 6.1. Data: The Data aspect connects input data and applications with enhancing trust in ML models.
- 6.1. Data: The surveyed methods include 196 model-agnostic or black-box techniques and 144 model-specific or white-box techniques.
6.6. Evaluation
The survey records the visual representations, ML task types, and interaction techniques used across the categorized literature, highlighting the breadth of visualization practice.
- Visual representation: 199 papers use bar charts, 115 use scatterplots or projections, and 86 use tables or lists.
- Machine learning types: 197 papers address supervised classification, while dimensionality reduction appears in 66 unsupervised-learning papers.
- Interaction technique: The most frequent interaction techniques are select, reconfigure, connect, and abstract or elaborate.
- Applications: The surveyed examples include image-focused visualization and systems for exploring data domains such as computer vision and humanities.
6.8. Target Group
The Target Group aspect distinguishes users by expertise and shows how visualization systems support beginners, practitioners, developers, and ML experts across tasks.
- 6.8. Target Group: The target groups include 196 beginners, 41 practitioners or domain experts, 162 developers, and 36 ML experts.
- Users and tasks: Visualization can help interpret and debug deep neural networks by presenting simpler models and training-process information.
- Users and tasks: Humanities experts can provide feedback that supervises clustering and can participate in user studies evaluating visualization effectiveness.
- Data and models: The survey distinguishes target variables such as binary, multi-class, multi-label, and continuous outcomes in classification and regression settings.
- Data and models: Visualization systems support model comparison, validation, debugging, and steering across supervised, unsupervised, and dimensionality-reduction tasks.
6.3. Machine Learning Processing Phase
The surveyed visualizations support multiple machine-learning processing phases through coordinated views, model-agnostic analysis, and diverse visual encodings. Across the 200 papers, techniques emphasize exploration, explanation, debugging, steering, and validation.
- Processing phases: Input, in-processing, and final phases correspond respectively to TL2, TL3, and TL5 trust categories.
- Model dependence: Model-agnostic techniques are twice as common as model-specific techniques, typically treating ML models as black boxes.ATMSeer exemplifies model-agnostic visualization by letting users steer AutoML search and explain results.
- Dimensionality: Nearly all visualizations are 2D: 196 papers use 2D displays, with one described exception using interactive 3D biplots.The 3D technique uses interactive bar-chart legends to show visible variables and help users choose viewing angles.
- Information representation: Visualizations use mapped or algorithmically derived information, including summary statistics, classifier feedback, aggregated information, and selected instances.ModelTracker uses computed summary statistics and charts, while Arendt et al.’s interface presents system-proposed instances for each class.
- Interaction and analysis: Multiple coordinated views support model overviews, training-process debugging, feature ranking, and exploration of embedding spaces.
- Visual encoding: Color is used in almost every surveyed paper, while opacity and size are the next most common visual variables.DeepCompare uses opacity and size to visualize deep-learning model results and support assessment of trade-offs.
6.6. Evaluation
Evaluation practices vary substantially across the surveyed visualization literature. Approximately half of the visualizations were never evaluated, while evaluated studies used domain experts, ML experts, or both, often through task-based usability assessment.
- Evaluation coverage: Approximately half of the surveyed papers did not include any type of evaluation.The report identifies evaluation as fundamental for validating visualization-tool and system usability.
- Evaluation methods: RuleMatrix evaluated usability by monitoring participants’ task accuracy and completion timing.The tool assists novice ML users in understanding, examining, and verifying predictive-model performance.
- User feedback: Five studies involved both domain and ML experts, 32 involved only domain experts, and 19 involved only ML experts.
- Evaluation coverage: One visualization tool was evaluated later in a new publication after its original paper omitted evaluation.
6.7. Trust Levels
The report categorizes trust in visualizations across five levels spanning data, learning methods, concrete models, and user expectations. The surveyed work addresses these levels through reliability, uncertainty, interpretability, collaboration, and evaluation-related techniques.
- Trust levels: The categorization identifies five trust levels: raw and processed data, learning method, concrete model, and user expectation.The passage names four grouped levels while describing the categorization as five levels overall.
- Data trust: Source reliability is associated with transparent collection processes and semantic exploration of erroneous data regions.AnchorViz lets users pin anchors, construct a topology over related instances, and examine discrepancies between semantically related points.
- Uncertainty and fairness: Visualization techniques represent prediction uncertainty and guide users toward potentially interesting parameter areas.One described approach uses 2D scatterplots and parallel coordinates for uncertainty investigation.
- Model trust: Interpretability and explainability are organized into understanding/explanation, debugging/diagnosis, refinement/steering, and comparison.These categories may occur in pairs or triplets within visualization systems.
- User expectation: Visualization supports collaboration, provenance, model comparison, and exploratory construction and validation of models.TensorFlow Graph Visualizer uses data-flow graphs and provenance, while LoVis supports progressive model construction and validation.
6.8. Target Group
Most visualization tools target domain experts and practitioners, with ML experts and developers also commonly included. Beginners and novice users are rarely considered as target groups.
- Primary audience: Most visualization tools cover domain experts or practitioners as a target group.
- Primary audience: ML experts and developers are commonly targeted together in addition to domain experts.
- Audience gaps: Beginners and novice users are rarely considered as target groups.RuleMatrix is an example of a tool designed to assist novice ML users.
7. Survey Data Analysis
The survey analysis identifies recurring visualization topics, prevalent application patterns, temporal and correlation structures, and underrepresented areas in research on trustworthy ML visualizations.
- Topic analysis: 13 papers in Topic 1 focus on hidden states and parameter spaces, especially time-series data and recurrent neural networks.Visualizing hidden states may preserve information that can enhance trust with appropriate expert intervention.
- Topic analysis: Topic 3 forms a tight cluster of five image-data papers focused mainly on deep learning hyper-parameters and reinforcement-learning rewards.The papers examine hyper-parameter selection and how high or low rewards develop during training.
- Topic analysis: 32 papers address models’ predictions through visualization with quality and validation metrics, including clustering challenges.The surveyed evaluations include participants ranging from novices to practitioners and ML experts.
- Topic analysis: Topic 7’s clustering and dimensionality-reduction work emphasizes distances, projections, and users’ cognitive expectations, forming a prominent 30-paper class.Topic 8 contains 18 papers focused on prediction-visualization prototypes and design choices informed by InfoVis research.
- Category analysis: 68 papers concern outlier detection, while computer vision, humanities, health, biology, supervised classification, dimensionality reduction, and clustering are common survey patterns.Post-processing and model-agnostic visualization techniques cover around 75% of all papers.
- Topic analysis: Topics 6 and 7 are most prominent and together cover approximately 35% of all papers, while shared terms explain several mixed clusters in the t-SNE embedding.The topic analysis also introduced subcategories that supported the paper’s broader categorization.
- Correlation analysis: The strongest negative correlation is between not-evaluated techniques and user-expectation evaluation, highlighting the need for further visualization evaluation.Positive correlations include stacking with boosting, deep-learning techniques with one another, and deep Q-networks with reinforcement learning.
8. Discussion and Research Opportunities
The discussion uses the survey browser and data-driven findings to identify trust-related gaps and research opportunities across bias, security, fairness, communication, and underrepresented ML categories.
- Survey browser: TrustMLVis Browser presents visualization-technique thumbnails with category-, time-, and text-based filtering and access to technique details and bibliographic information.The browser was developed as an interactive companion to the survey.
- Bias: Bias may arise from data equality, algorithm familiarity, model behavior, and user bias, while visualization systems can also scale poorly for massive data.The discussion identifies selection bias during dimensionality-reduction analysis as a specific concern.
- Bias: Combining interaction logs and analysis data with additional ML models is proposed as a possible way to guide improvements, but it requires empirical evaluation.The authors call for quantitative and qualitative experiments on combining automatic methods with smart visualizations.
- Alternatives and combination: Visualization research opportunities include comparing structures and algorithms, guiding data selection, detecting outliers, and comparing concrete model structures in situ.These methods are presented as possible remedies for trust compromises caused by data manipulation problems.
- Security vulnerabilities: Security research should address unethical attacks such as data poisoning from model, data-instance, feature, and local-structure perspectives.The discussion identifies visualization as a possible aid for avoiding adversarial vulnerabilities.
- Fairness and communication: Fairness, explainability, communication, and collaboration remain open challenges, including improving trust in visualizations themselves.User-friendly querying of specific instances and areas is identified as one starting point for communication.
- Almost unexplored areas: Underrepresented research areas include specific neural networks, boosting-only ensemble tools, and stacking ensemble learning.The discussion treats these categories as candidates for open challenges and novel research.
9. Conclusion
The STAR surveys 200 peer-reviewed publications on visualization for enhancing trust in ML models. It proposes a fine-grained categorization, analyzes topics and trends, and provides an interactive browser for exploring the findings.
- Conclusion: The survey covers 200 peer-reviewed publications and categorizes them into 8 high-level aspects, 18 category groups, and 119 categories.It also analyzes topic connections, category correlations, temporal trends, and data sets.
- Conclusion: TrustMLVis Browser makes the categorization and paper assignments publicly accessible through an interactive online exploration tool.The browser is intended to facilitate exploration of the survey’s information.