Source-linked AI summary
A Perspective on Explainable Artificial Intelligence Methods: SHAP and LIME
Ahmed Salih, Zahra Raisi-Estabragh, Ilaria Boscolo Galazzo, Petia Radeva, Steffen E. Petersen, Gloria Menegaz, Karim Lekadir
TL;DR
XAI methods seek to make black-box ML outputs more understandable, but the reliability of their explanations under model variation and feature collinearity remains a concern. This perspective compares how SHAP and LIME generate explanations and examines those issues through biomedical case studies. The authors report that both methods are highly affected by the adopted ML model and feature collinearity, warranting caution in their use and interpretation.
Problem
Evidence is limited on how model dependency and feature collinearity affect the reliability and interpretation of SHAP and LIME explanations.
Method
The paper compares SHAP and LIME and uses biomedical case studies to examine their explanation outputs, limitations, and possible solutions.
Results
SHAP and LIME outputs are highly affected by the adopted ML model and feature collinearity.
Takeaways & Limitations
Users should interpret SHAP and LIME outputs cautiously and understand their assumptions, including model dependency and feature independence.
Takeaways & Limitations
Both methods remain limited by feature dependence, nonlinear relationships, biased classifiers, uncertainty estimation, generalization, and inability to infer causality.
Abstract
from arXiv · showhide
eXplainable artificial intelligence (XAI) methods have emerged to convert the black box of machine learning (ML) models into a more digestible form. These methods help to communicate how the model works with the aim of making ML models more transparent and increasing the trust of end-users into their output. SHapley Additive exPlanations (SHAP) and Local Interpretable Model Agnostic Explanation (LIME) are two widely used XAI methods, particularly with tabular data. In this perspective piece, we discuss the way the explainability metrics of these two methods are generated and propose a framework for interpretation of their outputs, highlighting their weaknesses and strengths. Specifically, we discuss their outcomes in terms of model-dependency and in the presence of collinearity among the features, relying on a case study from the biomedical domain (classification of individuals with or without myocardial infarction). The results indicate that SHAP and LIME are highly affected by the adopted ML model and feature collinearity, raising a note of caution on their usage and interpretation.
1 Introduction
XAI emerged to make complex ML models more comprehensible by clarifying their decisions, influential features, and outcome certainty. This perspective examines how model dependency and feature collinearity affect XAI outputs, focusing on SHAP and LIME in a biomedical case study.
- XAI aims to demystify black-box ML models by making their operation and outputs more comprehensible to end-users.The motivation includes explaining specific decisions, influential features or regions, and the certainty of generated outcomes.
- Understanding model behavior is especially important for deploying advanced ML systems in high-risk fields such as healthcare.The paper links comprehensible explanations with reassurances needed for wider implementation.
- The perspective investigates how model dependency and feature collinearity affect the quality of XAI outcomes.It uses a biomedical case study and examines two common XAI methods, while also considering possible solutions to their limitations.
2 eXplainable artificial intelligence
SHAP and LIME are widely used XAI methods with different explanation scopes and computational approaches. The paper emphasizes that model dependency, collinearity, and nonlinear relationships can compromise their reliability and interpretation.
- Method comparison: SHAP provides local and global explanations, whereas LIME produces local explanations by fitting an interpretable surrogate model.SHAP considers feature coalitions, while LIME approximates a complex model locally with a linear model.
- Method comparison: SHAP can reflect nonlinear associations depending on the underlying model, while LIME's local linear surrogate cannot capture such associations.This distinction follows from the different mechanisms used to generate feature attributions.
- Limitations: Both methods are affected by feature collinearity, which can limit explanation reliability and undermine trust.SHAP may sample unrealistic instances under correlated features, while LIME treats features as independent.
- SHAP: SHAP outcomes depend on the selected ML model, so different classifiers can produce different explainability scores.The method is model-agnostic in applicability but model-dependent in its resulting explanations.
- Interpretation: SHAP scores should be interpreted through feature ranking rather than as direct feature weights, and biased classifiers can yield unrealistic explanations.The paper also notes limitations involving uncertainty estimates, generalization, feature dependence, and causal inference.
- LIME: LIME outcomes also depend on the selected model because the local surrogate explains that model's behavior for a specific instance.Its reported coefficients can omit nonlinear relationships present in the original model.
3 A case study
The case study examines how model choice and feature collinearity affect SHAP explanations. Different models produced different informative-feature lists, while NMR and MIP were used to assess and adjust robustness.
- Study design: The case study used 500 airline-passenger records, 22 features, and four classifiers: LGBM, LR, DT, and SVC.The data were divided into training and testing sets, with 20% used for testing.
- Model dependency: SHAP generated different lists of informative features for each model despite relatively similar accuracy, except for higher accuracy with LGBM.The authors caution that default model parameters prevent certainty that LGBM is superior.
- Model dependency: Two models identified Class as the most important feature, whereas the other two ranked different features first.This variation raises the question of which SHAP list should be trusted when features are collinear.
- Robustness assessment: NMR values were 0.231 for LGBM, 0.275 for LR, 0.445 for DT, and 0.273 for SVC.LGBM had the lowest NMR value, indicating the most robust corresponding outcome among these models.
- Collinearity adjustment: NMR identifies the more stable model but does not make SHAP account for feature collinearity; MIP modifies SHAP outputs to incorporate dependency.For LGBM, Ease of Online booking ranked fifteenth with SHAP and fifth after MIP.
4 Recommendations
The paper recommends presenting SHAP results with plots, plain-language explanations, and their assumptions, while comparing models and using NMR and MIP when features are collinear.
- Interpretation: Users should present SHAP results alongside plots, simple explanations, and assumptions such as feature independence and model dependence.The recommendation targets clearer interpretation of visualized SHAP outputs.
- Robustness: With collinear features, users should compare SHAP outcomes across different ML models to evaluate robustness when possible.The paper links this comparison to model-dependent informative-feature lists.
- Collinearity: NMR can help select the model producing the most stable informative-feature list, while MIP can enhance XAI outputs under collinearity.These are presented as post-hoc tools for assessing and modifying explanations.
5 Conclusions
The perspective discusses SHAP and LIME for tabular-data XAI and emphasizes that users, especially those without technical backgrounds, should understand their issues before applying them.
- Scope: The perspective discusses two widely used XAI methods, SHAP and LIME, particularly for tabular data.The paper frames these methods as tools whose outputs require informed interpretation.
- Practical implication: Users without technical backgrounds need awareness of the methods’ highlighted issues to apply them appropriately.The conclusion emphasizes appropriate use rather than treating explainability outputs as self-interpreting.
- Practical implication: The paper characterizes these considerations as significant and critical when XAI methods are implemented in any domain.This conclusion follows the perspective’s discussion of interpretation issues in XAI use.
6 Funding
The study acknowledges support from the British Heart Foundation, Fondazione CariVerona, and the National Institute for Health and Care Research.
- Funding: AMS was supported by a British Heart Foundation project grant, PG/21/10619.
- Funding: IBG and GM acknowledged support from Fondazione CariVerona through the EDIPO project.The passage identifies the 2018 research-science call and project reference 2018.0855.2019.
- Funding: ZRE’s Academic Clinical Lectureship post was supported by the NIHR Integrated Academic Training programme and a British Heart Foundation fellowship.The fellowship number is FS/17/81/ in the supplied passage.