Source-linked AI summary
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel
TL;DR
XAI explanation quality is difficult to assess reproducibly because existing approaches rely on costly human judgment or narrow proxy metrics. The paper introduces XAI-Arena, an LLM-as-a-judge framework for multidimensional, stakeholder-sensitive comparison of explanations. Human ratings show a strong aggregate association with LLM ratings, while relationships with technical proxy metrics depend on the property measured.
Problem
Existing XAI evaluation relies mainly on costly human judgment or technical proxy metrics, leaving scalable, reproducible, human-centered comparative assessment limited.
Method
XAI-Arena uses an LLM-as-a-judge framework to evaluate externally generated XAI explanations across multiple quality dimensions within a controlled protocol.
Results
Human ratings show a strong aggregate association with LLM ratings, whereas relationships with technical proxy metrics depend on the specific property measured.
Takeaways & Limitations
LLM-based evaluation can identify systematic differences in perceived XAI explanation quality across evaluation dimensions.
Takeaways & Limitations
The primary evaluation relies on a single LLM, so ratings may reflect model-specific tendencies.
Abstract
from arXiv · showhide
Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.
1 Introduction
XAI-Arena addresses the difficulty of systematically comparing XAI explanation quality by using LLMs as scalable, reproducible, multidimensional, and stakeholder-sensitive evaluators.
- Human evaluations provide rich insights but are costly to collect, difficult to scale, and variable across participant tasks or expertise.
- Proxy metrics scale evaluation but usually measure narrow technical properties and may not reflect how stakeholders perceive or use explanations.
- XAI explanations vary across method scope, output format, stakeholder expertise, and decision context, complicating comparative assessment.
- XAI-Arena uses a fixed-prompt, fixed-decoding LLM-as-a-judge protocol to assess explanations across eight quality dimensions and four stakeholder personas.
- The study compares SHAP, LIME, DiCE, partial dependence plots, and permutation importance across machine learning models, datasets, roles, and explanation formats.
- Human ratings are used to examine whether LLM ratings recover systematic differences and preference patterns rather than replace human assessment.
2 Related Work
Prior XAI evaluation combines human-centered studies and technical proxy metrics, but both approaches leave gaps in scalable, reproducible, and multidimensional assessment. XAI-Arena applies LLM-as-a-judge evaluation to explanations generated by external XAI methods.
- XAI Methods: XAI methods differ in explanatory scope, representation, abstraction level, and mechanism, including local attributions, counterfactuals, global plots, and feature importance.
- Evaluation of XAI Explanations: Existing comparative evaluations often focus on specific methods, tasks, or settings, limiting systematic assessment across models, datasets, formats, and stakeholders.
- Evaluation of XAI Explanations: The paper adopts eight complementary dimensions: perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and interpretability.
- Evaluation of XAI Explanations: Human-centered studies capture interpretation and trust but are expensive, time-consuming, and difficult to reproduce at scale.
- Evaluation of XAI Explanations: Proxy metrics quantify technical properties such as fidelity, completeness, stability, sensitivity, MoRF, and compactness without directly capturing human interpretation or usefulness.
- LLMs as Evaluators: LLM-as-a-judge research reports moderate to strong agreement with human judgments, but performance depends on prompt phrasing and the underlying model.
- LLMs as Evaluators: XAI-Arena fills the identified gap by evaluating explanations produced by external XAI methods rather than generating or verbalizing explanations itself.
3 The XAI-Arena Framework
XAI-Arena is a reproducible LLM-as-a-judge framework for evaluating XAI explanation artifacts across datasets, models, methods, formats, personas, and multiple quality dimensions. Its controlled workflow constructs standardized inputs, obtains ratings and justifications, and applies statistical analysis to compare explanation quality.
- Workflow: The framework follows six steps: model training, explanation generation, input construction, prompt assembly, LLM evaluation, and statistical analysis.The protocol is designed to support controlled and reproducible assessment across stakeholder roles and explanation formats.
- Evaluation output: The LLM evaluator outputs ratings and textual justifications for each evaluation dimension on a seven-point Likert scale.Ratings and justifications are stored with metadata such as dataset, model, explanation method, persona, format, and explanation identifier.
- Framework: XAI-Arena evaluates explanation artifacts from the perspectives of stakeholder personas, allowing the same artifact to receive role-dependent quality ratings.The evaluation input includes a persona representing roles such as a line manager or end user.
- Input and prompting: Each evaluation combines an explanation artifact with contextual information about the task, dataset, model, method, and, for local explanations, the prediction and selected feature values.A fixed prompt structure includes persona framing, textual context, the artifact, evaluation dimensions, a rating scale, and a constrained response format.
- Experimental instantiation: The instantiated benchmark combines synthetic and real-world datasets, multiple predictive models and XAI methods, stakeholder personas, and explanation formats.Synthetic datasets vary sample size, feature count, and linear versus nonlinear structure to isolate dataset effects on evaluation.
- Evaluation corpus: The final corpus contains 17,252 structured LLM evaluation records and 138,016 individual dimension ratings.The records include ratings and textual justifications across eight dimensions; 19 records produced incorrect output formats.
4 Results
The results show that LLM ratings distinguish related but distinct dimensions of explanation quality and vary systematically with datasets, models, explanation methods, and stakeholder personas. Human validation finds strong aggregate alignment between LLM and human ratings.
- Dimension-level patterns: Clarity receives the highest mean rating, while transparency and actionability receive the lowest, indicating higher ratings for understandability than for revealing model logic or supporting action.Perceived simplicity and interpretability follow clarity among the highest-rated dimensions.
- Dimension-level patterns: Clarity, perceived simplicity, and interpretability are strongly correlated, whereas actionability correlates more weakly with the other dimensions.Faithfulness and transparency are also closely associated, and the dimensions capture related but distinct aspects of explanation quality.
- Dataset effects: Ratings are broadly similar across synthetic and real-world datasets, with synthetic datasets receiving slightly higher ratings on most dimensions and real-world datasets higher ratings for task adequacy and transparency.Faithfulness has the largest dataset-origin difference, while overall differences remain modest and rating patterns are largely stable.
- Model effects: Model differences are strongest for transparency and faithfulness, with logistic regression rated highest on both and model effects significant across all eight dimensions.The largest effects are transparency (F=1362.40, ω2=.283) and faithfulness (F=215.93, ω2=.056).
- Persona effects: Stakeholder personas significantly affect all eight dimensions: end users generally give the lowest ratings, while ML developers and data scientists tend to give higher ratings.The largest persona effects occur for perceived simplicity (ω2=.711), interpretability (ω2=.630), and clarity (ω2=.596).
- Proxy metrics: Stability metrics align strongly with LLM-rated faithfulness, led by rank stability (ρ=.749), cosine similarity (ρ=.738), and top-k overlap (ρ=.664), while classification and regression correlations are weakly negative.MoRF AUC shows a very weak negative correlation with faithfulness (ρ=−.127, p=.007).
- Human validation: Human and LLM ratings show a strong positive association across 32 paired observations (Spearman’s ρ=.693, p<.001), with higher human ratings generally matched by higher LLM ratings.The mean absolute error is 0.866, RMSE is 1.003, and the mean LLM–human difference is 0.255.
5 Discussion
The discussion finds that XAI explanation quality varies across dimensions, datasets, models, methods, and stakeholders, supporting multidimensional and stakeholder-sensitive evaluation. XAI-Arena offers a scalable complement to human evaluation, but its evidence remains bounded by study scope and evaluator dependence.
- Interpretation: LLM ratings distinguish meaningful aspects of explanation quality, with related dimensions following coherent but non-identical patterns.The dimensions therefore capture related but distinct aspects of explanation quality.
- Dataset and model sensitivity: Dataset and model choice affect explanation-quality ratings, but not all dimensions equally; model differences are most pronounced for transparency and, to a lesser extent, faithfulness.The observed pattern is broadly consistent with linear models being easier to interpret than more complex models such as neural networks.
- Method-specific strengths and weaknesses: No XAI method leads across every evaluation dimension, so method selection should depend on the explanation goal rather than an overall ranking.SHAP receives strong ratings for faithfulness and transparency but does not consistently lead on more user-centered dimensions.
- Stakeholder sensitivity: End Users receive the lowest ratings for perceived simplicity, clarity, interpretability, and actionability, underscoring the importance of stakeholder perspective.Explanations rated favorably from a technical perspective may still be unsuitable for intended users.
- Validation: LLM-rated faithfulness aligns strongly with three stability metrics but only weakly negatively with MoRF AUC, because proxies capture specific technical properties.Human-rated explanations also tend to receive higher LLM ratings, although strong relative association does not imply exact score agreement.
- Validation: LLM ratings provide a useful aggregate signal, but human judgments, LLM evaluations, and proxy metrics remain complementary rather than interchangeable.XAI-Arena is intended to support large-scale screening and identify cases requiring focused human investigation, not replace human evaluation.
- Framework implications: XAI-Arena supports systematic, multidimensional, stakeholder-sensitive comparisons by making method-specific strengths and weaknesses visible.Its protocol uses standardized prompts, fixed dimensions, documented artifacts, and consistent experimental settings.
- Limitations: The study is limited to tabular data and selected models, XAI methods, proxy metrics, and mainly text-only comparisons.Ratings may also depend on prompts, personas, dimension definitions, repeated calls, and model-specific tendencies; broader prompt and cross-run testing is needed.
6 Conclusion
The conclusion reports that LLM-based evaluation identifies systematic differences in perceived XAI explanation quality across dimensions. XAI-Arena combines standardized ratings, multiple quality dimensions, and stakeholder perspectives into a scalable and reproducible comparative framework, while human ratings show strong aggregate association with LLM ratings.
- Conclusion: LLM-based evaluation identifies systematic differences in perceived XAI explanation quality across evaluation dimensions.The conclusion presents XAI-Arena as the first LLM-as-a-judge framework for multidimensional, human-centered XAI explanation evaluation.
- Conclusion: Human ratings show a strong aggregate association with LLM ratings, whereas technical-proxy relationships depend on the specific property measured.This indicates converging but property-dependent validation evidence.
- Conclusion: XAI-Arena combines standardized LLM ratings with multiple quality dimensions and stakeholder perspectives for scalable, reproducible comparative assessment.The framework is positioned as a structured way to compare XAI explanations.
AI Use Disclosure
Generative AI tools assisted with research-code implementation and debugging, while the authors controlled experimental design, code integration, verification, and result interpretation.
- AI use disclosure: Generative AI tools assisted with implementation and debugging of the research code.The disclosure distinguishes coding assistance from author-controlled research decisions.
- AI use disclosure: The authors performed and controlled all experimental design, code integration, verification, and interpretation of results.
A Evaluation Protocol
XAI-Arena evaluates generated explanations through fixed, persona-specific prompts that combine explanation artifacts with task and model context. The protocol logs ratings and justifications across eight dimensions, with fixed seeds, version-controlled code, and validation checks supporting reproducibility.
- Protocol: The six-step protocol trains models, generates local and global XAI artifacts, constructs contextual inputs, evaluates them with an LLM, and analyzes logged outputs.Local methods include SHAP, LIME, and DiCE; global methods include PDP and permutation importance.
- Protocol: Each evaluation input combines an explanation artifact with textual context describing the prediction task, model setting, and explanation scope.Local contexts include the prediction f(x) and selected feature values; global contexts describe the dataset–model setting.
- Reproducibility: Ratings and textual justifications are stored with dataset, model, method, persona, instance, and format metadata, then aggregated for statistics and correlations.Experiments use fixed random seeds and version-controlled code.
- Validation: Dataset, model, and artifact validation checks structural consistency, numerical integrity, metadata alignment, and required files before downstream evaluation.The pipeline also checks class distributions or regression-target distributions and validates stored model bundles.
D.3 Detailed Statistical Results for Model Effects
The model-effects analysis finds statistically significant differences across all eight evaluation dimensions, with the strongest effects for transparency and faithfulness. Logistic regression generally receives higher ratings than nonlinear models, while MLP receives the lowest ratings in several comparisons.
- Overall model effects: All eight evaluation dimensions show statistically significant predictive-model effects, with the largest effect for transparency (𝜔2 = .283).Faithfulness has the next-largest effect (𝜔2 = .056), while effects for the remaining dimensions are small.
- Faithfulness: Logistic regression receives higher faithfulness ratings than MLP, random forest, and XGBoost.The reported differences are ΔM = 0.81, ΔM = 0.59, and ΔM = 0.60, respectively, all with p < .001.
- Transparency: Logistic regression receives higher transparency ratings than MLP, while model comparisons show substantial variation in transparency.The section reports a large logistic-regression advantage over MLP, with ΔM = 1.57 and g = 2.13.
- Nonlinear models: Among nonlinear models, MLP receives lower ratings than random forest and XGBoost, while random forest also rates below XGBoost.All reported pairwise differences are statistically significant at p < .001.
- Artifact validation: Only artifacts passing structural and method-specific validation checks are included in downstream evaluation.Remaining warnings reflect approximation- or method-related behavior rather than invalid artifact generation.
E.8 Games–Howell Comparisons for XAI Method Effects
XAI methods exhibit distinct evaluation profiles across dimensions: SHAP leads faithfulness and transparency, DiCE leads actionability, and permutation importance leads interpretability. Text+plot explanations receive higher ratings than text-only explanations across all dimensions.
- Faithfulness: SHAP receives the highest faithfulness rating, exceeding permutation importance, PDP, DiCE, and LIME.The largest reported difference is versus LIME (ΔM = 2.23, g = 3.41), and LIME receives the lowest faithfulness rating.
- Transparency: SHAP receives the highest transparency rating, whereas DiCE receives the lowest.SHAP exceeds DiCE by ΔM = 1.68 with g = 2.53.
- Interpretability: Permutation importance receives higher interpretability ratings than SHAP and DiCE.The reported differences are ΔM = 0.35 and ΔM = 0.31, respectively.
- Explanation format: Text+plot explanations receive consistently higher ratings than text-only explanations across all eight dimensions (p < .001 throughout).Ratings across the two formats remain strongly positively correlated for LIME, SHAP, permutation importance, and PDP.
G Pairwise Comparisons between Stakeholder Personas
Stakeholder persona substantially shapes explanation ratings, with End Users generally assigning lower scores than technical or managerial personas. Differences are largest for perceived simplicity, clarity, task adequacy, trust calibration, actionability, and interpretability, but smaller for transparency and faithfulness.
- Perceived simplicity: Perceived-simplicity ratings differ most strongly by persona, with ML Developer, Data Scientist, and Manager ratings exceeding End User ratings.The reported differences are ΔM = 1.84, ΔM = 1.75, and ΔM = 1.34, respectively.
- Clarity: ML Developer, Data Scientist, and Manager ratings exceed End User ratings for clarity by ΔM = 1.21, ΔM = 1.10, and ΔM = 1.00.All reported pairwise differences are statistically significant at p < .001.
- Trust calibration: For trust calibration, ML Developer ratings exceed Data Scientist and End User ratings by ΔM = 0.74 and ΔM = 0.62.Manager ratings also exceed End User ratings by ΔM = 0.39.
- Actionability: For actionability, ML Developer, Manager, and Data Scientist ratings exceed End User ratings by ΔM = 1.05, ΔM = 0.83, and ΔM = 0.55.The corresponding effect sizes are g = 1.66, g = 1.46, and g = 1.03.
- Transparency and faithfulness: Persona differences are smallest for transparency and faithfulness, with no significant differences among the three non-End-User personas for transparency.For faithfulness, ML Developer and Manager ratings do not differ significantly (p = .912).
H Cross-LLM Robustness
The evaluator-robustness analysis replayed 9,988 text-only evaluation prompts with two alternative LLMs and compared their ratings with GPT-5.4 across dimensions. The resulting correlations were generally positive, indicating broadly similar rating patterns across evaluator models.
- Evaluation setup: 9,988 text-only evaluation prompts were replayed using Claude Opus 4.7 and Mistral Small 4.These alternative evaluators were used to assess dependence on evaluator-model choice.
- Evaluation setup: Ratings from GPT-5.4 were compared with those from the alternative evaluators using Spearman rank correlations for each evaluation dimension.The correlations were computed across matched text-only evaluation instances.
- Results: The cross-LLM results generally showed weak positive or strong positive correlations across dimensions.Table 10 reports the pairwise correlations between evaluator models.
- Results: The broadly positive correlations imply that different LLM evaluators capture broadly similar patterns in explanation-quality ratings.This supports cross-LLM robustness of the ratings across the evaluated dimensions.