Source-linked AI summary
Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
Harini Suresh, Steven R. Gomez, Kevin K. Nam, Arvind Satyanarayan
TL;DR
Black-box systems require explanations that are understandable, relevant, and useful to diverse stakeholders, yet prior frameworks often organize stakeholders primarily by expertise or role. The paper introduces a framework that separates knowledge from needs and evaluates its ability to describe, assess, and generate research opportunities. It finds richer stakeholder–need intersections, more precise user-focused evaluations, and opportunities for reflexive interpretability research.
Problem
Interpretability research lacks sufficiently precise ways to characterize diverse stakeholders and their needs, limiting explanations’ relevance to people beyond system builders.
Method
The framework represents formal, instrumental, and personal knowledge across ML, data-domain, and milieu contexts, and organizes needs into goals, objectives, and tasks.
Results
Coding 58 papers found that the framework consistently describes stakeholder knowledge and needs while adding granularity, supporting precise evaluations, and generating new intersections for study.
Takeaways & Limitations
The framework supports more precise participant recruiting and comparative user studies while enabling a more reflexive approach to interpretability research.
Takeaways & Limitations
The framework is not exhaustive and remains a living artifact that may need new goals, objectives, tasks, and finer-grained knowledge or milieu categories as evidence accumulates.
Abstract
from arXiv · showhide
To ensure accountability and mitigate harm, it is critical that diverse stakeholders can interrogate black-box automated systems and find information that is understandable, relevant, and useful to them. In this paper, we eschew prior expertise- and role-based categorizations of interpretability stakeholders in favor of a more granular framework that decouples stakeholders' knowledge from their interpretability needs. We characterize stakeholders by their formal, instrumental, and personal knowledge and how it manifests in the contexts of machine learning, the data domain, and the general milieu. We additionally distill a hierarchical typology of stakeholder needs that distinguishes higher-level domain goals from lower-level interpretability tasks. In assessing the descriptive, evaluative, and generative powers of our framework, we find our more nuanced treatment of stakeholders reveals gaps and opportunities in the interpretability literature, adds precision to the design and comparison of user studies, and facilitates a more reflexive approach to conducting this research.
1 INTRODUCTION
Interpretability mechanisms must serve diverse stakeholders with different goals, but existing approaches often fail to identify those users precisely. The paper responds with a granular framework that separates stakeholder knowledge from interpretability needs.
- Black-box ML systems are difficult to understand, limiting effective trust-building and accountability for decisions they affect.
- Stakeholders in medical ML—including physicians, patients, medical staff, and developers—need different kinds of information about model outputs and behavior.
- Interpretability methods often omit intended users, making explanations understandable mainly to developers or insufficiently useful to people in practice.
- Prior frameworks commonly categorize stakeholders by expertise or functional role, then derive needs from those categories.
- The proposed framework combines formal, instrumental, and personal knowledge across ML, data-domain, and milieu contexts with a three-level hierarchy of goals, objectives, and tasks.
- Coding 58 papers showed that the framework describes stakeholder knowledge and needs with greater granularity, supports more precise evaluations, and generates new intersections for study.
2 BACKGROUND AND MOTIVATION
Interpretability research has increasingly studied users, but expertise- and role-based frameworks inadequately represent the composable, changing relationships among stakeholder attributes, roles, and needs.
- Interpretability research seeks human-oriented definitions because many techniques neither specify target users nor identify the tasks explanations should support.
- Expertise-based frameworks classify users along roughly linear scales and infer needs such as education for novices or debugging and deployment for experts.
- Role-based frameworks distinguish stakeholders such as creators, operators, executors, examiners, executives, engineers, and end users.
- Role categories can conflate expertise and needs: auditors may want instance-level insight without equivalent ML expertise, while doctors and patients may share a role but require different explanations.
- Role-based accounts also treat roles as relatively static even though users’ roles and interpretability needs may change with exposure and familiarity.
- Expertise-based accounts can overlook lived experience, tacit knowledge, and the effects of everyday familiarity with technology.
- Interpretability goals cut across expertise and roles, including trust, bias assessment, contesting decisions, recourse, and domain-specific testing.
- Existing frameworks lack a sufficiently granular and composable vocabulary for distinguishing stakeholder attributes, roles, and interpretability needs.
3 A FRAMEWORK TO CHARACTERIZE THE STAKEHOLDERS OF INTERPRETABLE ML
The framework characterizes interpretability stakeholders through separate knowledge types and contexts, then independently organizes their needs into goals, objectives, and tasks. This decoupling exposes needs that cut across conventional roles and supports more precise, context-sensitive interpretability design.
- Framework development: The framework was developed through iterative literature review and cross-disciplinary divergence and convergence to create a granular, composable stakeholder vocabulary.The process extracted descriptions of stakeholders, needs, actions, and goals, then drew on domains outside interpretability.
- Knowledge and context: Stakeholder expertise is decomposed into formal, instrumental, and personal knowledge rather than treated as a single scale.Formal knowledge concerns codified theories, instrumental knowledge concerns applying them through tools, and personal knowledge includes experience, memories, and shared values.
- Knowledge and context: These knowledge types are situated across machine learning, the data domain, and milieus such as the physical, social, or institutional environments of human-AI interaction.The contexts determine which knowledge is relevant to researching, deploying, interpreting, and responding to ML systems.
- Knowledge and context: Decoupling knowledge from context helps distinguish stakeholders hidden by broad categories such as model users or model breakers and suggests different interface requirements.The framework allows designers to consider, for example, toolkit-oriented interfaces for ML instrumentalists and designs that support personal knowledge or example-based explanation.
- Stakeholder needs: Interpretability needs are treated independently from stakeholder expertise and roles, allowing goals such as debugging or contesting decisions to span diverse stakeholders.Personal knowledge may help identify model errors, while lawyers, judges, and activists may contest decisions to affect systematic change.
- Stakeholder needs: Needs form a hierarchy of long-term goals, short-term objectives, and immediate tasks with many-to-many relationships and latent temporal ordering.The chronology is represented through broad, iterative ML-process phases rather than a precise linear sequence, and tasks are omitted from the chronology to preserve their many-to-many mapping.
4 EVALUATION & EXAMPLE APPLICATIONS OF THE FRAMEWORK
The framework describes prior interpretability work with finer-grained stakeholder and need categories, then supports evaluation, gap analysis, and generation of new design possibilities. Its applications also connect interpretability design to participant selection, personal knowledge, and sociotechnical power.
- Descriptive Power: The framework described stakeholders and needs across 58 interpretability papers, with all knowledge-context intersections and goal/objective/task categories appearing in multiple papers.The most frequent knowledge categories were ML-Instrumental (23/58), Data Domain-Formal (19/58), and Data Domain-Instrumental (17/58).
- Descriptive Power: The most common objectives were explaining model-influenced decisions and debugging or improving models, while understanding influential factors and detecting mistaken behavior were the most common tasks.Both O4 and O1 appeared in 21/58 papers; T4 appeared in 11/58 and T2 in 8/58.
- Descriptive Power: The framework adds granularity by distinguishing stakeholders previously grouped as “lay users,” and connects concepts that different papers describe with different terminology.Its composable representation also exposes many-to-many relationships among knowledge contexts, goals, objectives, and tasks.
- Implications: The framework reveals underrepresentation of needs such as understanding data use and contesting model-based decisions, alongside limited attention to formal milieu knowledge.It also identifies personal knowledge as less observable through traditional evaluation and development processes, making elicitation an underexplored design opportunity.
- Evaluative Power: The framework supports more ecologically valid and appropriately scoped evaluations by refining participant-pool definitions and identifying acceptable forms of proxying.Residents or medical students may share relevant formal medical knowledge with doctors, whereas loan-explanation studies may require similar personal milieu knowledge.
- Generative Power: Decoupling stakeholder knowledge from needs enables new persona-need combinations, including non-ML experts who seek model debugging or improvement.The framework also supports considering underexplored stakeholders, societal power dynamics, and affected people’s input when designing interpretability systems.
5 LIMITATIONS AND FUTURE WORK
The framework is useful but intentionally incomplete: it exposes gaps in interpretability research while leaving room to expand its treatment of knowledge and expertise. Future work should extend the framework beyond epistemological accounts to include rhetorical construction.
- Findings: The framework covers 58 papers and reveals gaps in the interpretability literature while offering a richer intersection of stakeholder expertise and needs.It also supports more precise comparison and design of user-focused evaluations and a more reflexive research process.
- Scope: The authors describe the framework as a living artifact rather than an exhaustive account of stakeholder goals, objectives, and tasks.They expect it to grow as machine learning is deployed more deeply in existing domains.
- Knowledge model: The knowledge model compresses additional forms of expertise, including informal, contingent, tacit, meta-, and cultural knowledge, into broader categories.The authors identify this reduced granularity as an opportunity for future refinement.
- Rhetorical expertise: The expertise model is grounded only in epistemology, although expertise is also constructed rhetorically through an audience’s granting of expert status.Future work could adapt findings from data visualization on authority, uncertainty, framing, trust, and recall to interpretability.