Source-linked AI summary
A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations
Ana Naveriani, Jakob Suchan, Stefano Zoia, Mehul Bhatt, Antonio Lieto, Gian Luca Pozzato
TL;DR
Metaphor explanation evaluation has focused mainly on holistic ratings, offering limited insight into the structure of quality and disagreement. This paper introduces a six-dimensional cognitively motivated framework and finds that explanation quality has dominant underlying dimensions, while automatic evaluation partially recovers this structure.
Problem
Existing metaphor explanation evaluation relies mainly on holistic ratings, providing limited insight into the conceptual mappings and meanings explanations should convey.
Method
The paper proposes six complementary cognitive dimensions and studies them through dense ratings of 100 explanations by 16 annotators, totaling 11,200 ratings.
Results
The six dimensions organize into a correlated Mapping–Emergence–Purpose cluster, two near-independent dimensions, and an automatic pipeline that predicts Emergence best (R2 = 0.76).
Takeaways & Limitations
Multidimensional evaluation provides diagnostic insight into metaphor explanation quality and supports automatic evaluators that preserve human judgment structure.
Takeaways & Limitations
The framework was evaluated only on a deliberately constructed benchmark, and the automatic evaluation study used a relatively small dataset for feasibility rather than state-of-the-art claims.
Abstract
from arXiv · showhide
Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.
1 Introduction
Metaphor explanations are difficult because they must make conceptual mappings, emergent meanings, familiar experiences, and interpretive context explicit. The paper therefore introduces a cognitively motivated multidimensional evaluation framework and examines its reliability, redundancy, disagreement, and automatic recoverability.
- Metaphors help people understand abstract concepts through more familiar and concrete experiences.
- Explaining metaphorical meaning is considerably harder than intuitively understanding metaphorical expressions.
- Effective explanations must reveal conceptual mappings, account for emergent meaning, connect figurative language to familiar experiences, and provide interpretive context.
- The paper introduces a cognitively motivated multidimensional framework and analyzes its reliability, redundancy, and disagreement.
- An exploratory automatic evaluation study investigates whether standard language encoders can recover these evaluative aspects.
2 A Multidimensional Framework for Evaluating Metaphor Explanations
The paper proposes a cognitively motivated multidimensional framework for evaluating metaphor explanations. Drawing on established theories of metaphor, analogy, comprehension, and cognitively inspired computational approaches, the framework comprises six complementary dimensions.
- Framework foundations: The framework comprises six dimensions for evaluating metaphor explanations.These dimensions capture complementary aspects considered in assessment.
- Framework foundations: The framework builds on Conceptual Metaphor Theory, analogical reasoning, psycholinguistic models of metaphor comprehension, and cognitively inspired computational approaches.The cited foundations include work by Lakoff and Johnson, Gentner, Glucksberg, Bowdle and Gentner, Lieto et al., and Cappa et al.
3 Analysis of Human Judgments
A dense annotation study shows that human judgments of metaphor explanations are systematically structured across the proposed dimensions, with substantial variation in agreement and clear correlations among criteria. Overall quality is most strongly predicted by Emergence and Purpose, while disagreement peaks for borderline explanations.
- Study Design: 11,200 individual human quality ratings covered 100 metaphor explanations, 16 annotators, six dimensions, and an overall score on a five-point Likert scale.The corpus comprised 100 explanations, each independently rated across all criteria.
- Inter-rater Reliability: Agreement varied by dimension: Emergence had the highest Krippendorff’s α (α = 0.579), while Non-Circularity had the lowest (α = 0.370).No dimension reached the conventional threshold for “reliable” agreement (α0.8).
- Correlational Structure: Mapping, Emergence, and Purpose formed a strongly correlated cluster (ρ = 0.76–0.89, p < 10−25), while Non-Circularity and Origin & Cultural Accuracy were near-independent (|ρ| ≤0.21).Accessibility correlated moderately with the cluster (ρ ≈0.47– 0.54), indicating that the six dimensions do not function as equally distinct criteria.
- Overall Quality: Emergence was the strongest independent predictor of Overall quality (β = 0.499, p < .001), followed by Purpose (β = 0.290, p < .001).The estimates came from a linear model including all six standardized dimensions simultaneously.
- Disagreement: Disagreement followed a U-shaped relationship with explanation quality, peaking for borderline cases and falling for explanations receiving consistently high or low ratings.“Love is a Fine Wine” received near-unanimous ratings, whereas “A Dead End” produced the largest variation across annotators.
4 Preliminary Automatic Evaluation
A standard BERT-based pipeline partially recovers the framework’s multidimensional and human judgment structure, especially for Emergence and Purpose. However, this feasibility study does not establish competitive automatic evaluation, and model error’s association with human disagreement is not significant.
- Pipeline: The pipeline fine-tunes BERT-base to predict six dimensions, combines predictions with a gradient-boosted judge model, and uses metaphor-level grouped k-fold cross-validation.No metaphor appears in both training and test folds within a fold.
- Prediction performance: R2 = 0.76 for Emergence, 0.68 for Purpose, and 0.51 for Mapping; performance is lower for Accessibility, Non-Circularity, and Origin & Cultural Accuracy.The reported lower R2 values are 0.34, 0.19, and 0.13, respectively.
- Feasibility and limitations: MAE= 0.69 and R2 = 0.37 for overall quality, while model error correlates positively but non-significantly with human disagreement: Pearson r = 0.397, p = .083.Spearman ρ = 0.203, p = .392, and 13 of 20 predictions (65%) fall within one human standard deviation of the mean rating.
- Learned structure: The judge model weights Emergence and Purpose most heavily, mirroring human analysis, while their correlation is ρ = 0.89 and their prediction R2 values are 0.76 and 0.68.This suggests the evaluator captures the dominant structure underlying human judgments while assigning comparatively little weight to the remaining dimensions.
- Feasibility and limitations: The study is a feasibility check and does not claim competitive automatic evaluation.This limitation qualifies the preliminary findings.
5 Discussion and Outlook
The paper presents a cognitively motivated multidimensional framework for evaluating metaphor explanations and empirically characterizes six complementary cognitive constructs. An exploratory automatic evaluation pipeline partially recovers this structure, supporting the diagnostic value of multidimensional evaluation.
- The framework evaluates metaphor explanations through six complementary cognitive constructs.
- The empirical study reveals distinct patterns of reliability, redundancy, and disagreement across the dimensions.
- An exploratory automatic evaluation pipeline partially recovered the structure identified by multidimensional evaluation.
Limitations
The framework was evaluated only on a deliberately constructed benchmark, limiting its demonstrated applicability to naturally occurring explanations. The annotation study also does not establish that the proposed dimensions are exhaustive or universal.
- Evaluation scope: The framework has been evaluated only on a deliberately constructed benchmark, not on naturally occurring explanations from humans or contemporary language models.This limits the evidence for how well the framework generalizes beyond the benchmark setting.
- Framework coverage: The annotation study shows that the dimensions capture meaningful aspects of human evaluation but does not establish that they are exhaustive or universal.The framework is theory-informed, but its completeness and generality remain unconfirmed.
Ethical Statement
The study evaluates metaphor explanations without deploying language models in real-world decision-making. It uses informed consent, anonymous ratings, and a benchmark free of personal or sensitive information.
- Ethical Statement: The work does not deploy language models in real-world decision-making.It focuses on evaluating metaphor explanations.
- Ethical Statement: Participants provided informed consent and submitted anonymous ratings of metaphor explanations.
- Ethical Statement: The benchmark contains linguistic examples and intentionally constructed explanation failures without personal or sensitive information.
Supplementary Material · A Additional Human Analyses · A.1 Annotator Disagreement Across Quality Levels
The supplementary material extends the paper’s analyses, with A.1 showing that annotator disagreement varies systematically across Overall-quality levels. Disagreement is lowest for clearly poor or successful explanations and highest for intermediate ones.
- Supplementary Material: The appendix provides supplementary analyses supporting the main paper’s results.It includes disagreement analyses, regression diagnostics, automatic-evaluation analyses, participant characteristics, and reproducibility details.
- A Additional Human Analyses: Additional human analyses examine annotator disagreement across Overall-quality levels.This analysis is presented in subsection A.1 of the supplementary material.
- A.1 Annotator Disagreement Across Quality Levels: Figure 2 summarizes disagreement across buckets of mean Overall quality.The figure reports average annotator disagreement across these quality buckets.
- A.1 Annotator Disagreement Across Quality Levels: Disagreement is operationalized as the standard deviation of Overall ratings assigned to each metaphor.This measure captures variation among annotators’ ratings for the same metaphor.
- A.1 Annotator Disagreement Across Quality Levels: The disagreement pattern is approximately U-shaped across mean Overall-quality buckets.Agreement is therefore not uniform across the quality scale.
- A.1 Annotator Disagreement Across Quality Levels: Annotators show greater agreement when explanations are judged clearly poor or clearly successful.These quality extremes correspond to lower disagreement.
- A.1 Annotator Disagreement Across Quality Levels: Disagreement is highest for explanations receiving intermediate scores.Intermediate-quality explanations produce the strongest divergence in annotator judgments.
A.2 Disagreement Across Evaluation Dimensions … B.1 Criterion-Level Generalization
The paper finds systematic variation in annotator disagreement across evaluation dimensions and metaphor examples, while regression and automatic-evaluation analyses reveal uneven criterion predictability and generalization. These results identify which dimensions most strongly relate to Overall quality and which are harder to predict on held-out metaphors.
- A.2 Disagreement Across Evaluation Dimensions: Purpose shows the greatest average annotator variability, followed by Mapping, whereas Origin shows the lowest variability.This suggests stronger convergence on a metaphor’s source or cultural basis than on its communicative purpose or conceptual mapping.
- A.3 Examples of High and Low Annotator Agreement: Annotators converge on explanations that are clearly successful or clearly unsuccessful, while highest disagreement generally occurs for explanations with intermediate mean scores.The pattern is consistent with the aggregate relationship between disagreement and Overall quality.
- A.4 Full Regression Results: Emergence provides the strongest independent predictive signal for Overall quality, followed by Purpose, while Accessibility lacks a significant independent association after controlling for the other dimensions.The regression estimates associations with Overall quality while holding the remaining dimensions constant.
- B Additional Analyses of the Automatic-Evaluation: The automatic-evaluation analysis shows uneven criterion-level performance, with held-out prediction errors generally exceeding training errors across dimensions.This pattern indicates some overfitting in the modest-sized dataset.
- B.1 Criterion-Level Generalization: Purpose and Mapping exhibit the largest increases in held-out error, indicating comparatively greater difficulty generalizing these dimensions.The comparison uses mean absolute error on training and held-out test sets.
- B.1 Criterion-Level Generalization: Non-Circularity shows slightly lower error on the held-out test set than on the training set.This is the exception to the general pattern of higher held-out error across most dimensions.
B.2 Association Between Human Variability and Model Error … D.1 Annotation Protocol
The exploratory analyses found no reliable statistical association between human disagreement and automatic-evaluation error, while documenting participant diversity and an independent, asynchronous annotation protocol. Together, these results qualify the variability findings and describe how the annotations were collected.
- B.2 Association Between Human Variability and Model Error: r = .397 (p = .083) for Pearson correlation and ρ = .203 (p = .392) for Spearman correlation indicated a positive but uncertain relationship between human disagreement and evaluator error.The evidence did not establish a reliable association between human and model difficulty.
- B.3 High- and Low-Disagreement Groups: σ = 0.935 defined the median split separating held-out metaphors into high- and low-disagreement groups based on Overall-rating variability.Error was descriptively greater for high-disagreement examples, but the comparison was exploratory because few held-out metaphors were available.
- B.3 High- and Low-Disagreement Groups: U = 58.00 (p = .571) showed that high- and low-disagreement groups did not differ significantly under the Mann–Whitney test.The group comparison concerned automatic-evaluation error across the median-split disagreement groups.
- B.3 High- and Low-Disagreement Groups: t = 1.07 (p = .303) likewise showed no statistically significant high- versus low-disagreement difference under Welch’s t-test.This second test reinforced the exploratory, non-significant nature of the group comparison.
- C Participant Information: Participants represented diverse linguistic, cultural, geographic, educational, and disciplinary backgrounds, with information recorded on language, English proficiency, familiarity, age or education, and recruitment.Recorded familiarity included linguistics, literature, and metaphor.
- D.1 Annotation Protocol: The annotation task was administered through a Google Forms survey presenting the objective, 1–5 rubric, and six evaluation criteria.The survey header contained the task instructions and criteria shown in Figure 4.
- D.1 Annotation Protocol: Annotations were completed independently and asynchronously without supervision or participant communication, then exported from Google Forms as a CSV for analysis.Participants did not complete the task in a single supervised session.
D.2 Explanation Generation and Benchmark Construction … D.5 Reproducibility Checklist
The study constructs a 100-explanation metaphor benchmark with controlled criterion-level failures and evaluates it using a two-stage automatic pipeline. Training uses metaphor-level splits to prevent leakage, with reproducibility details including a held-out set of 20 metaphors and single-run reporting.
- D.2 Explanation Generation and Benchmark Construction: The benchmark contains 100 metaphors spanning physical, social, cultural, and abstract domains, with candidate explanations generated by GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro.A standardized prompt produced approximately 25-word explanations in a broadly uniform, professional tone.
- D.2 Explanation Generation and Benchmark Construction: The six criteria assess Non-Circularity, Origin and Cultural Accuracy, Emergence, Accessibility, Purpose, and Mapping.Reference explanations were required to pass all six criteria, while targeted and mixed failures were designed to violate specified subsets.
- D.2 Explanation Generation and Benchmark Construction: The final benchmark comprises 15 reference explanations, 60 targeted failures, and 25 mixed failures.Targeted failures include ten explanations violating each individual criterion, and outputs were manually filtered and regenerated when they were hallucinated, off target, or inconsistent with the requested pattern.
- D.3 Automatic Evaluation Pipeline: The automatic evaluator first predicts six criterion scores and overall human judgment, then maps the criterion predictions to a final overall quality estimate.Stage 1 concatenates the metaphor and explanation and uses bert-base-uncased with a seven-output regression head; Stage 2 uses a Gradient Boosting regressor.
- D.3 Automatic Evaluation Pipeline: Stage 2 uses five-fold GroupKFold cross-validation grouped by metaphor to prevent leakage, with 200 boosting estimators.The Gradient Boosting regressor was implemented in scikit-learn and trained on the training portion.
- D.4 Training Configuration: Stage 1 was trained on 1,600 annotation instances from 16 annotators and 100 metaphor explanations, with metaphor-level splitting and 20 held-out metaphors.The hyperparameters were manually set using standard defaults for BERT fine-tuning on a small regression dataset.
- D.5 Reproducibility Checklist: The reproducibility checklist specifies a 100-explanation dataset, six criterion targets plus overall quality, metaphor-level splitting, bert-base-uncased, and a scikit-learn Gradient Boosting regressor.Primary software included PyTorch, Hugging Face Transformers, and scikit-learn; hardware used an Apple M2 with PyTorch MPS and CPU fallback where required.
- D.5 Reproducibility Checklist: Reported results use one train–test split and one training run rather than averages across random seeds.This reporting choice is explicitly listed in the reproducibility checklist.