Source-linked AI summary

Manipulating and Measuring Model Interpretability

Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, Hanna Wallach

arXiv:1802.07810v5cs.AIcs.CY

TL;DR

Machine-learning interpretability research has proposed many supposedly interpretable models, but has relatively little experimental evidence about their behavioral effects. This paper runs pre-registered experiments varying feature count and transparency in functionally identical models. Clear, few-feature models improved simulation but not beneficial reliance, and clear models impaired detection of sizable mistakes, underscoring the need to test interpretability rather than rely on intuition.

  • Problem

    Relatively few experimental studies examine whether supposedly interpretable models make people follow beneficial predictions or detect model mistakes.

  • Method

    N=3,800 participants completed pre-registered experiments varying model feature count and transparency while measuring simulation, reliance, and mistake correction.

  • Results

    Clear models with few features improved prediction simulation, but did not increase beneficial reliance and reduced detection and correction of sizable mistakes.

  • Takeaways & Limitations

    Interpretability factors can have negligible or detrimental effects on behavior, so developing interpretable models requires empirical testing rather than intuition.

  • Takeaways & Limitations

    The experiments studied laypeople using linear regression for real-estate valuation, so other stakeholders, models, tasks, and domains may yield different findings.

Abstract

from arXiv · show

With machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Although many supposedly interpretable models have been proposed, there have been relatively few experimental studies investigating whether these models achieve their intended effects, such as making people more closely follow a model's predictions when it is beneficial for them to do so or enabling them to detect when a model has made a mistake. We present a sequence of pre-registered experiments (N=3,800) in which we showed participants functionally identical models that varied only in two factors commonly thought to make machine learning models more or less interpretable: the number of features and the transparency of the model (i.e., whether the model internals are clear or black box). Predictably, participants who saw a clear model with few features could better simulate the model's predictions. However, we did not find that participants more closely followed its predictions. Furthermore, showing participants a clear model meant that they were less able to detect and correct for the model's sizable mistakes, seemingly due to information overload. These counterintuitive findings emphasize the importance of testing over intuition when developing interpretable models.

1 INTRODUCTION

The paper treats interpretability as a latent, human property that should be studied by manipulating model factors and measuring behavioral outcomes. Pre-registered experiments found that clearer, simpler models improved simulation but did not reliably improve reliance or error detection.

  • Motivation: Interpretability lacks consensus in definition and measurement, with simplicity, transparency, simulatability, and trustworthiness often conflated.The paper argues that different stakeholders may require different forms of interpretability.
  • Conceptual framework: Interpretability is framed as a latent human property influenced by manipulable model factors and reflected in measurable behavior.Relevant outcomes include simulating predictions, following beneficial predictions, and detecting model mistakes.
  • Experiments: N=3,800 participants completed pre-registered experiments varying feature count and model transparency while measuring simulation, reliance, and mistake correction.The models were studied in real-estate valuation with lay participants.
  • Findings: Participants better simulated clear two-feature models, but did not follow their predictions more closely when doing so was beneficial.This comparison was made against relevant alternative model presentations.
  • Findings: Clear models reduced participants’ ability to detect and correct sizable mistakes, while the clearest eight-feature condition produced especially poor behavioral outcomes.The authors associate the latter pattern with information overload.
  • Implications: The findings emphasize testing interpretability interventions rather than relying on intuition about clear model internals.The paper reports that commonly presumed interpretability factors can have negligible or detrimental behavioral effects.

2 RELATED WORK

Related work spans interpretability techniques, human-centered research on mental models and sensemaking, and decision-making studies of trust in computational aids. These traditions motivate studying how people understand and use models.

  • Interpretability research: Interpretability research has developed techniques for explaining complex models, but relatively few studies test how these factors affect people’s behavior.The paper positions its controlled experiments within this empirical gap.
  • Human–computer interaction: Human–computer interaction research treats people as active participants who form mental models of computational systems.Prior work also examines intelligibility, transparency, explanations, and user modification.
  • Sensemaking: Sensemaking research studies how people organize information and build situation awareness when interacting with people or computational systems.In this paper’s context, it concerns understanding machine-learning models and their data.
  • Decision making: Decision-making research examines people’s aversion to or trust in computational aids and ways to increase that trust.Related work also studies simple or improper linear models resembling those used here.

3 EXPERIMENT 1: PREDICTING APARTMENT SELLING PRICES

Experiment 1 tested whether model feature count and transparency affected simulation, reliance, and mistake correction in apartment-price predictions. Clear, two-feature models improved simulation but did not increase beneficial reliance and reduced correction of sizable mistakes.

  • H1. Simulation: Participants with clear, two-feature models had lower simulation errors than participants in the other primary conditions.This result supported H1 and was statistically significant, t(994) = −12.06, p < 0.001.
  • H2. Deviation: Participants did not follow clear, two-feature model predictions more closely than black-box, eight-feature model predictions when beneficial.The comparison was not significant, contradicting H2.
  • H3. Detection of mistakes: Participants in the primary model conditions predicted higher prices than baseline participants for apartments with unusual configurations.The models made overly high predictions for these apartments, so larger deviations could indicate mistake correction.
  • H2. Deviation: Using a clear model reduced participants’ deviations from model predictions, contrary to the expectation that they would independently correct sizable mistakes.Across the four primary conditions, deviations differed significantly, with clear-model participants deviating less than black-box-model participants.
  • H3. Detection of mistakes: Participants shown clear models were less accurate than those shown black-box models when predicting apartment 11’s selling price.The transparency main effect was significant, F(1, 994) = 31.98, p < 0.001.
  • Prediction errors: Participants’ prediction errors did not differ significantly across the four primary conditions, although using a model was advantageous overall for typical apartments.For typical configurations, model-assisted participants had lower prediction errors than the no-model baseline, but would have benefited from following the model more closely.

4 EXPERIMENT 2: REPRESENTATIVE U.S. PRICES

Experiment 2 replicated the first experiment using representative U.S. prices. Clear two-feature models improved simulation but did not increase beneficial reliance, and clear models impaired mistake detection on sufficiently unusual apartments.

  • 4 EXPERIMENT 2: REPRESENTATIVE U.S. PRICES: The second experiment replicated the first experiment’s main findings after scaling prices and fees to representative U.S. levels.The design otherwise remained unchanged from Experiment 1.
  • 4.2 Findings: Participants shown a clear, two-feature model had lower simulation errors than participants in the other primary conditions.This result was significant, t(594) = −10.41, p < 0.001.
  • 4.2 Findings: There was no significant difference in beneficial model-following between the clear two-feature and black-box eight-feature conditions.The comparison was t(594) = 0.49, p = 0.626.
  • 4.2 Findings: For apartment 12, clear-model participants followed the model’s overly high prediction more closely, producing worse final predictions.The condition comparison was significant, t(594) = −4.16, p < 0.001; apartment 11 showed no significant clear-versus-black-box difference.
  • 4.2 Findings: $4,000, roughly 3% of the $120,000 average selling price, was the maximum pairwise prediction-error difference across primary conditions.The overall ANOVA was significant, F(3, 594) = 8.60, p < 0.001.

5 EXPERIMENT 3: WEIGHT OF ADVICE

Experiment 3 tested whether an alternative advice-taking measure and a “Human Expert” label changed reliance on model predictions. Neither weight of advice nor final-prediction distance showed greater reliance on clear models, and the human-expert label did not increase reliance.

  • 5 EXPERIMENT 3: WEIGHT OF ADVICE: Experiment 3 used weight of advice alongside absolute prediction distance to measure how closely participants followed model predictions.Weight of advice captures updating toward the model’s prediction from an initial prediction.
  • 5.2 Findings: The clear two-feature model did not produce significantly closer beneficial following than the black-box eight-feature model.This null result held for both weight of advice and absolute distance from the model’s prediction on typical apartments.
  • 5.2 Findings: Participants followed predictions labeled “Human Expert” no more closely than predictions from black-box models.The authors suspect increasing experience with the source label during the experiment may explain the difference from prior findings.
  • 5.2 Findings: Unlike Experiments 1 and 2, clear-model participants were not less able to detect and correct overly high predictions for apartments 11 or 12.This result motivated the fourth experiment.

6 EXPERIMENT 4: OUTLIER FOCUS AND DETECTION OF MISTAKES

Experiment 4 removed a potential anchoring route and tested whether an outlier-focus message could mitigate clear models’ disadvantage on unusual apartments. The message improved mistake correction and eliminated the clear-versus-black-box difference.

  • 6 EXPERIMENT 4: OUTLIER FOCUS AND DETECTION OF MISTAKES: Experiment 4 removed the simulation step and varied whether participants saw an outlier-focus message highlighting unusual apartments.The experiment was designed to test anchoring and information-overload explanations.
  • 6.1 Findings: Participants shown an outlier-focus message deviated more from model predictions for apartment 6 and apartment 8.Both comparisons were significant: t(791) = −4.72, p < 0.001, and t(795) = −5.00, p < 0.001.
  • 6.1 Findings: Without an outlier-focus message, clear-model participants deviated less than black-box participants for both unusual apartments.The comparisons were significant for apartment 6, t(393) = −3.65, p < 0.001, and apartment 8, t(395) = −3.51, p < 0.001.
  • 6.1 Findings: With an outlier-focus message, clear- and black-box-model participants did not differ significantly in deviations for either apartment.For apartment 6, t(401) = −0.004, p = 0.996; for apartment 8, t(394) = −1.64, p = 0.101.
  • 6.1 Findings: The findings support information overload as an explanation for earlier clear-model disadvantages and identify outlier focus as a mitigation.The proposed overload concerns visual attention to unusual configurations rather than working-memory cognitive load.

7 LIMITATIONS

The experiments have several scope and design limitations, including a narrow stakeholder, model, and domain sample, constrained model comparisons, possible feature-specific confounds, and limited process measurement.

  • Scope: The experiments studied laypeople using linear regression for real-estate valuation, so findings may differ for other stakeholders, tasks, models, or domains.Suggested extensions include data scientists and domain experts, classification, decision trees, rule lists, deep neural networks, and domains such as medical diagnosis or hiring.
  • Design constraints: The first three experiments forced two-feature and eight-feature models to make the same predictions, excluding settings where complex deep models outperform simpler ones.This constraint avoided confounding presentation effects with model fidelity and large feature-count differences.
  • Possible confounds: Participants may have reacted to particular feature combinations or to having information unavailable to the model, potentially influencing judgments about the models.The authors also note that a negative coefficient for total rooms may have been confusing or viewed as incorrect.
  • Measurement: The experiments lacked process measures and were short and one-shot, limiting insight into participants’ cognitive and sensemaking processes.Interviews, think-aloud protocols, process tracing, and longitudinal measurement could provide deeper explanations of behavior.

8 DISCUSSION AND CONCLUSION

The experiments produced counterintuitive findings: clearer, simpler models did not increase beneficial reliance and could impair mistake detection. The authors therefore recommend testing interpretability goals behaviorally and considering presentation strategies that reduce information overload.

  • Findings: Participants did not significantly follow clear models with few features more than black-box models with more features, despite lower prediction errors when simply following the model.The result challenges the expected link between model clarity, simplicity, and beneficial reliance.
  • Findings: Clear models hampered participants’ ability to detect sizable mistakes, seemingly because the amount of detail caused information overload.An outlier-focus message eliminated this behavior in a post-hoc investigation.
  • Findings: The clear, eight-feature condition produced the worst simulation, reliance, and selling-price prediction accuracy among the primary conditions.This pattern was consistent with the idea that too much information can be detrimental.
  • Implications: Alerting people to possible outliers and eliciting their predictions before revealing the model may encourage closer inspection of individual cases.The authors propose an auxiliary model for outlier detection and asking for users’ predictions first.
  • Conclusion: The number of features and model transparency should not be ignored, but interpretability goals should be assessed through testing rather than intuition.The authors emphasize that different interpretability goals may require evaluating different behavioral outcomes.

APPENDICES

The appendices document decision-aid examples, experiment instructions, apartment selection procedures, and tables describing model inputs, apartment configurations, predictions, and errors.

  • Decision-aid examples: Examples cover decision aids for malignancy risk, company success, educational outcomes, wildfire risk, hospital readmission, bail, pre-trial release, and venture success.The examples specify model features and, in some cases, information available to users but not used by the model.
  • Experiment instructions: The study instructions ask participants to predict New York City apartment prices with a model during training and testing phases.Participants observe model predictions and actual prices during training, then make predictions for twelve new apartments during testing.
  • Model instructions: In the clear two-feature condition, the model uses bathrooms and square footage, applying feature weights and subtracting a $260,000 adjustment factor.Each bathroom contributes $350,000, and the square-footage contribution is added before the adjustment is subtracted.
  • Participant tasks: Testing requires participants to predict the model’s output, assess confidence in that prediction, and then estimate the actual selling price and confidence.The training phase similarly asks participants to estimate actual prices from model predictions and then review outcomes.
  • Stimulus construction: The apartment-selection procedure matched rounded predictions across two- and eight-feature models and sampled apartments to represent different prediction errors and configurations.Training and testing assignments were randomized subject to error and configuration constraints.
  • Appendix tables: Tables report apartment configurations and model predictions, prediction errors, and error fractions for training and testing stimuli.The tables cover experiments 1–3 and specify the different apartment subsets used in experiment 4.

Appendix D EXPERIMENT 3 HYPOTHESES AND FINDINGS

Experiment 3 preregistered hypotheses concerned deviation from model predictions, weight of advice, comparisons with human experts, and detection of model mistakes.

  • Hypotheses: H7 predicted that participants would deviate less from a clear model with few features than from a black-box model with many features.This hypothesis targets behavioral agreement with the model’s predictions.
  • Hypotheses: H8 predicted higher weight of advice for participants viewing a clear model than for those viewing a black-box model with many features.The hypothesis concerns how strongly participants incorporate model advice.
  • Hypotheses: H9 predicted that deviation and weight-of-advice measures would differ between black-box model predictions and human-expert predictions.The comparison varies the source of the advice while retaining the same behavioral measures.
  • Hypotheses: H10 predicted condition differences in participants’ ability to correct inaccurate model predictions on unusual examples.The hypothesis extends the mistake-detection question to experiment 3 conditions.

D.1 Results

The results report no significant differences in deviation, weight of advice, or mistake correction across key model-presentation conditions, despite differences in prediction simulation.

  • There was no significant difference in participants’ deviation from the model between clear-2 and bb-8.The reported test was t(798) = −0.87, p=0.384.
  • There was no significant difference in participants’ weight of advice between clear-2 and bb-8.The reported test was t(819) = 1.27, p=0.205.
  • There was no significant difference in deviation or weight of advice between bb-8 and the expert condition.The reported tests were t(994) = 0.45, p=0.655 for deviation and t(1005) = −0.38, p=0.704 for weight of advice.
  • Participants in the clear conditions were no less able to correct inaccurate predictions for apartments 11 and 12.The reported contrasts were not significant for either apartment.

F.1 Experiment 1: Predicting Prices

Experiment 1 analyzed simulation error, prediction deviation, and deviation for two apartments using two-way ANOVA results.

  • Experiment 1 reports a two-way ANOVA on simulation error.
  • Experiment 1 reports a two-way ANOVA on deviation between the model’s and participants’ price predictions.
  • Experiment 1 separately reports deviation analyses for apartment 11 and apartment 12.

F.2 Experiment 2: Scaled-down prices

Experiment 2 analyzed simulation error and prediction deviation, including separate analyses for apartments 11 and 12, using two-way ANOVA.

  • Experiment 2 reports a two-way ANOVA on simulation error.
  • Experiment 2 reports a two-way ANOVA on deviation between the model’s and participants’ price predictions.
  • Experiment 2 separately reports deviation analyses for apartment 11 and apartment 12.

F.3 Experiment 3: Weight of Advice

Experiment 3 analyzed deviation and weight of advice across its four primary conditions and all conditions, including a human-expert condition.

  • Experiment 3 reports two-way ANOVA results on deviation across the four primary conditions.
  • Experiment 3 reports one-way ANOVA results on deviation across all conditions, including the human-expert condition.
  • Experiment 3 reports two-way ANOVA results on weight of advice across the four primary conditions.
  • Experiment 3 reports one-way ANOVA results on weight of advice across all conditions, including the human-expert condition.
  • Experiment 3 separately reports deviation analyses for apartment 11 and apartment 12.
Loading 1802.07810v5…