Source-linked AI summary
The State of the Art in Integrating Machine Learning into Visual Analytics
A. Endert, W. Ribarsky, C. Turkay, W Wong, I. Nabney, I Díaz Blanco, Fabrice Rossi
TL;DR
As data becomes more complex, people need analytical systems that support trustworthy and interpretable reasoning. This state-of-the-art report surveys how machine learning has been integrated into visual analytics, synthesizes selected advances, and identifies opportunities and challenges for deeper integration. The surveyed work includes integrated systems and emerging approaches that place interaction alongside visualization and machine learning, while also exposing limits in current cognitive and user-process models.
Problem
Growing data scale and complexity make it harder for people to reason about data while maintaining valid, trustworthy, and interpretable conclusions.
Method
The paper provides a comprehensive survey of machine-learning methods and visual analytics systems, then organizes opportunities for further integration between the two areas.
Results
The survey finds that existing systems integrate machine learning with visualization and exploration, while interactive machine learning can be embedded in the human-computer process rather than used only as static preprocessing.
Takeaways & Limitations
Future research can formalize steerable machine learning, improve coupled interaction and visualization, and better determine how analytical tasks and effort should be divided between humans and machines.
Takeaways & Limitations
A detailed computational model of users’ analytical processes remains constrained because existing task models are too high-level and a more detailed task analysis is needed.
Abstract
from arXiv · showhide
Visual analytics systems combine machine learning or other analytic techniques with interactive data visualization to promote sensemaking and analytical reasoning. It is through such techniques that people can make sense of large, complex data. While progress has been made, the tactful combination of machine learning and data visualization is still under-explored. This state-of-the-art report presents a summary of the progress that has been made by highlighting and synthesizing select research advances. Further, it presents opportunities and challenges to enhance the synergy between machine learning and visual analytics for impactful future research directions.
1. Introduction
As data grows in scale and complexity, people need tools that support valid, trustworthy, and interpretable reasoning. This report surveys progress integrating machine learning with visual analytics and identifies opportunities for tighter coupling between user activities and computational models.
- Increasing data scale and complexity make reasoning more difficult, increasing the need for powerful tools that preserve trustworthy and interpretable results.
- Visual analytics and machine learning have complementary strengths and weaknesses for helping people make sense of data.
- Existing visual analytics systems use selected machine-learning models, but broader integration could couple models to user tasks and activities.
- The report summarizes advances at the intersection of machine learning and visual analytics and highlights future research opportunities in both disciplines.
2. Models and Frameworks
The report frames visual analytics through models of human sensemaking, visualization processes, interaction, and machine learning. Together, these models expose how computation and human reasoning can be connected, while also revealing limitations and opportunities for more interactive integration.
- Models of Sensemaking and Knowledge Discovery: Sensemaking models describe the cognitive processes people use to organize data, form explanations, and gain insight.They provide a human-centered basis for designing systems that support analytical reasoning.
- Models of Sensemaking and Knowledge Discovery: Pirolli and Card’s model separates sensemaking into foraging, where information is gathered, and synthesis, where users construct and test hypotheses.Existing visualization tools often focus on only one of these phases.
- Models of Sensemaking and Knowledge Discovery: The Data-Frame and knowledge-generation models represent sensemaking as exchanges among data, human knowledge, visualization, models, actions, and findings.The KGS model explicitly incorporates Prior Knowledge and User Knowledge into an iterative knowledge-generation process.
- Models of Sensemaking and Knowledge Discovery: These frameworks clarify relationships among machine learning, visualization, exploration, and knowledge generation, but no current visual analytics system embodies all components of the Sacha and KGS models.Existing systems such as VAiRoma may use machine learning as static preprocessing, whereas the KGS model illustrates a role for interactive ML.
- Models of Interactivity in Visual Analytics: Semantic interaction infers users’ analytical reasoning from visual interactions and uses it to steer underlying analytic models during exploration.This approach binds model steering to interactive affordances and supports co-reasoning between people and analytic models.
- Machine Learning Models and Frameworks: Interactive machine learning models use repeated user feedback to add training examples and improve a classifier’s approximation of the classified phenomenon.The report also notes that formal user inputs may not match the actions people naturally take during exploratory data analysis.
3. Categorization of Machine Learning Techniques Currently used in Visual Analytics
The report categorizes machine-learning integration in visual analytics by algorithm type and by users’ analytical intent. It surveys how interactive systems modify parameters or computation domains and how they incorporate analytical expectations, drawing on literature from both visualization and machine-learning venues.
- The taxonomy organizes visual-analytics literature by algorithm type—dimension reduction, clustering, classification, and regression/correlation analysis—and by interaction intent.The interaction-intent dimension focuses on how analysts try to improve machine-learning results.
- Dimension reduction distills high-dimensional information so conventional visualizations can be used and important features identified, while clustering identifies groups of similar instances.The report notes that these tasks frequently require both computation and user expertise.
- Modify Parameters & Computation Domain: Modify parameters and computation domain lets users adjust algorithm parameters, computational measures, algorithms, or the data subset to which computation applies.Interactive visual representations support selecting subsets of data and variables before running algorithms.
- Define Analytical Expectations: Define analytical expectations involves communicating expected results to computational methods, including user knowledge in interactive learning processes.In unsupervised learning, this form of user input falls within the broader semi-supervised framework.
4. Application Domains
Visual analytics integrates machine learning with interactive visualization across text, financial, multimedia, and streaming-data applications. These systems support organizing complex data, exploring relationships, externalizing insights, and building knowledge, while multimedia integration and continuously updated data remain challenging.
- Text Corpora: Text visual analytics uses computational methods such as topic modeling, clustering, and dimension reduction to organize and explore large unstructured collections.Interactive systems can connect topics with time, location, and people, while visual encodings represent document similarity and support information foraging.
- Text Corpora: Spatial workspaces support synthesis by letting analysts manually arrange documents into clusters and externalize insights during sensemaking.Large, high-resolution displays can promote spatially oriented analysis.
- Text Corpora: VAiRoma combines topic, time, and geographic analysis to narrate 3,000 years of Roman and Italian history from 189,000 Wikipedia articles.Its linked views expose geographic hotspots, article lists, and event peaks for selected topics and periods.
- Multimedia Visual Analytics: Multimedia visual analytics faces an open challenge in aligning image or video features with text, because automated approaches are error-prone and often require user guidance.Semantic concepts and relationships must be maintained across data types.
- Streaming Data: Financial visual analytics combines risk or preference models with interactive views, enabling users to inspect overall and subtle sector- and country-level market changes.A market visualization covers assets in 3 countries and 28 sectors from 2006 to 2009 and marks the 2008 stock-market crash.
- Streaming Data: Social-media visual analytics combines sentiment analysis, statistical and machine-learning models, and interactive exploration to support movie box-office prediction.In use cases involving 4 films, several non-expert participants outperformed experts under the authors’ criteria; supervised learning still requires a training dataset.
5. Embedding Steerable ML Algorithms into Visual Analytics
The report argues for steerable machine-learning algorithms embedded directly within visual analytics, so users can manipulate models and data while observing evolving visual outcomes. It distinguishes response regimes and presents dynamically evolving models as a foundation for interactive steering, while noting that this area remains largely unexplored.
- Motivation: Steerable ML binds interaction to visualization and machine learning, extending interactive ML toward deeper analytic reasoning and knowledge feedback loops.The report frames this integration as a route toward richer sensemaking and actionable knowledge.
- Steerable Dimension Reduction: Interactive dimension reduction lets users modify algorithm parameters or input data and inspect intermediate results during iteration.Iteration-level visualization exposes evolving outputs and supports real-time interaction with those results.
- Dynamic-System Formulation: In the dynamic-system formulation, input data x and parameters w form context u, internal state y evolves under f, and output v represents the resulting visualization.Changing x or w drives the algorithm toward a new steady state and a new visualization; continuous f can yield smooth animated transitions.
- Steering Context: User steering can incorporate prior feature relevance by modifying metric weights, including diagonal elements of the weight matrix in interactive PCA and distance-based systems.Users can alter dimension contributions or item distances and recompute projections based on their perceived similarities.
- Interactive Response Levels: Real-time responses are defined as < 0.1 second, while direct manipulation spans 0.1 to 2-3 seconds for richer visual-analytic interactions.VAiRoma’s 2-3-second geographic update did not appear to hinder reasoning during interface evaluation.
- Interactive Response Levels: Delays up to 2-3 seconds, and perhaps longer, may be digestible for some interactive ML algorithms rather than requiring real-time response.The report presents this as a potential way to reduce the response burden, while calling for further user studies.
6. Open Challenges and Opportunities for ML and VA
Open challenges center on making ML and VA mutually steerable, interpretable, and responsive to analysts’ reasoning. The report identifies opportunities to learn from user interaction, expose computation during convergence, balance human and machine effort, and support both exploratory and rigorous analysis.
- Creating and Training Models from User Interaction Data: User interaction logs can steer computation beyond explicit labeling by capturing analysts’ interests and exploratory behavior.Broader interaction-aware systems can express users’ mental models, preferences, and expertise while keeping them engaged in exploration.
- Creating and Training Models from User Interaction Data: Learning from interaction can reveal new analytic tasks and motivate ML algorithms that model how people analyze data.The report frames this as a continued science of interaction linking discoveries about human analytic processes to advances in ML.
- Balancing Human and Machine Effort, Responsibility, and Tasks: Generalizable empirical evidence is still needed to determine how users and machines should divide or co-complete the many subtasks in analysis.The report also proposes additional metrics for evaluating balance of effort in mixed-initiative systems.
- Enhancing Trust and Interpretability: Combined ML and VA tools should connect creative, tentative exploration with critical inquiry that produces deliberate and rigorous explanations.The report names this design principle “fluidity and rigour” and encourages concurrent visualization of qualitative data and computed features.
- Steerable Machine Learning: Rendering intermediate ML results could let users tune training data, parameters, or cost functions during algorithmic convergence.This direction calls for ML algorithms designed around new interaction mechanisms rather than exposing only final solutions.
- Beyond Current Methods: Modeling analysts’ processes could support expert strategies and less-trained users, but current process models require more detailed task analysis.The report notes that high-level frameworks do not yet provide enough detail for a full computational model of analytic activity.
7. Conclusions
The paper surveys machine learning methods and visual analytics systems, then identifies opportunities for deeper integration. It calls for steerable machine learning, richer coupled interaction, improved human–machine task division, and collaboration across disciplines.
- The paper surveys machine learning methods and visual analytics systems that integrate machine learning.
- Future systems could combine steerable machine learning with visualization and interaction that provide more advanced user feedback.
- The paper highlights dynamically dividing tasks between humans and machines and developing metrics for balancing their effort.
- The authors argue that collaboration between machine learning and visual analytics researchers can support more impactful systems for gaining insight into data.
9. Author Bios
The authors work across visual analytics, machine learning, interactive visualization, and data analysis. Their research spans foundational leadership, interpretable systems, human–computer interaction, and applications to complex processes and domains.
- Alex Endert directs the Visual Analytics Lab at Georgia Tech, while William Ribarsky directs the Charlotte Visualization Center at UNC Charlotte.Ribarsky is also identified as a founder of visual analytics and a former VAST Steering Committee chair.
- Cagatay Turkay develops methods that combine interactive visualizations and computational tools for informed analysis processes.William Wong studies visual analytics sense-making in high-information-density domains and among low-literacy users.
- Ian Nabney researches machine learning, data visualization, time series, and Bayesian methods, with applications including condition monitoring, biomedical engineering, and urban science.Ignacio Díaz Blanco researches intelligent data analysis, visualization, control, and signal processing for understanding and optimizing complex systems.
- Fabrice Rossi researches machine learning and data analysis, with particular interest in interpretable systems.