Source-linked AI summary

Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks

Annalisa Szymanski, Noah Ziems, Heather A. Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, Ronald A. Metoyer

arXiv:2410.20266v1cs.HC

TL;DR

The paper examines how LLM-as-a-Judge evaluation compares with expert evaluation in tasks where evaluation is difficult because many answers may be correct. Its findings identify strengths and weaknesses in LLM–SME evaluations and emphasize keeping SMEs involved in evaluation systems.

  • Problem

    Evaluation is difficult in tasks with an extremely large set of possibly correct answers, especially where access to trained-professional expertise is limited.

  • Method

    The paper compares the LLM-as-a-Judge evaluation approach with evaluations involving subject matter experts.

  • Results

    The evaluations reveal strengths and weaknesses in judgments produced by LLMs and SMEs.

  • Takeaways & Limitations

    Evaluation systems for tasks requiring expert knowledge should keep SMEs in the loop.

  • Takeaways & Limitations

    The work does not thoroughly explore the correlation between pre-training data content and downstream performance on domain-specific tasks.

Abstract

from arXiv · show

The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous research indicates that LLM-as-a-judge exhibits a strong correlation with human judges in the context of general instruction following. However, for instructions that require specialized knowledge, the validity of using LLMs as judges remains uncertain. In our study, we applied a mixed-methods approach, conducting pairwise comparisons in which both subject matter experts (SMEs) and LLMs evaluated outputs from domain-specific tasks. We focused on two distinct fields: dietetics, with registered dietitian experts, and mental health, with clinical psychologist experts. Our results showed that SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health when evaluating overall preference. Additionally, the results indicated variations in SME-LLM agreement across domain-specific aspect questions. Our findings emphasize the importance of keeping human experts in the evaluation process, as LLMs alone may not provide the depth of understanding required for complex, knowledge specific tasks. We also explore the implications of LLM evaluations across different domains and discuss how these insights can inform the design of evaluation workflows that ensure better alignment between human experts and LLMs in interactive systems.

1 INTRODUCTION

Existing evaluation methods struggle with open-ended, domain-specific tasks, motivating comparison of LLM judges with subject matter experts. The study examines this alignment in dietetics and mental health using mixed-methods pairwise evaluations.

  • Motivation: Automated overlap metrics become less useful when open-ended tasks permit multiple valid responses.BERTScore also may miss user preferences, domain-specific nuances, and expert reasoning requirements.
  • Motivation: Human evaluation can assess open-ended outputs but is expensive, slow, and preference-inconsistent.
  • Motivation: LLM judges offer inexpensive and reproducible evaluation, but positional, knowledge, and format biases may remain incompletely eliminated.
  • Research gap: Domain-specific complex tasks require evaluation that accounts for analytical reasoning, intuition, pattern recognition, and practical expertise.
  • Study aims: The study compares LLM and SME judgments for complex tasks in dietetics and mental health through pairwise evaluations and explanation analysis.The research questions address agreement and factors contributing to evaluation differences.
  • Contributions: The paper argues for incorporating SMEs into evaluation processes and tailoring expert input to domain and task differences.Its contributions include empirical evidence, explanation analysis, agreement variability, and integration implications.

2 BACKGROUND

Open-ended and specialized outputs are difficult to evaluate because they may have many valid answers and require judgments beyond token or semantic overlap. LLM judges address evaluation cost and reproducibility, but their alignment with SMEs in complex domains remains uncertain.

  • Evaluation challenges: Token-overlap metrics can fail when a correct generated answer paraphrases the labeled answer.Semantic metrics such as BERTScore measure similarity but do not resolve all evaluation requirements.
  • Evaluation challenges: Open-ended generations can contain many valid answers, making conventional labeled-answer evaluation difficult.Some generations may mix correct and incorrect sentences, further complicating assessment.
  • Human evaluation: Human pairwise evaluation compares two generations for the same instruction but is expensive, time-consuming, and difficult to reproduce.
  • Related approaches: Prior work includes systems that compare prompts, expose evaluation reasons, or align grading with human-defined criteria.
  • LLM judges: LLM evaluators have shown high correlation with human preferences, yet their alignment with SMEs in specialized or complex domains remains uncertain.

3 METHODS

The study focuses on expert knowledge in dietetics and mental health, where evaluation requires clinical judgment and domain-specific understanding. It uses curated instructions, aspect questions, and pairwise comparisons between LLM and SME judgments, including an expert-persona condition.

  • Domain selection: The study selects dietetics and mental health because their decisions rely on clinical judgment, evidence-based practice, and practical expertise.
  • Domain selection: LLM biases and reliance on training-data patterns may cause difficulty assessing expert-level decisions and overlooking critical factors.
  • Dataset construction: The researchers created 25 domain-specific instructions designed to challenge model accuracy, complex responses, and clinical judgment.
  • Dataset construction: The instructions covered dietetics themes including disease management, dietary preferences, lifestyle, nutrition, health, allergies, and sensitivities.
  • Dataset construction: Mental health instructions were based on conversational agents, emotional support, online counseling, and topics including stress, anxiety, depression, suicide, and PTSD.
  • Evaluation design: The pairwise method presents LLMs and SMEs with the same two candidate outputs and compares their overall and aspect-specific preferences.The study also tests expert personas and uses explanations to examine evaluation differences.

4 EXPERIMENT

The experiment compares LLM and expert judgments on the same domain-specific candidate outputs. It recruits dietitians and clinical psychologists, collects overall and aspect preferences with explanations, and uses AlpacaEval-based LLM judging with randomized response order.

  • Participants: Ten registered dietitians and ten clinical psychologists completed domain-tailored evaluation surveys.Participants were recruited for education and experience in their respective fields.
  • Participants: Dietitians averaged 9.6 years of experience and psychologists averaged 15.2 years.The participant profiles also recorded AI use and familiarity.
  • Models: Candidate responses were generated with GPT-4o and GPT-3.5-turbo, while GPT-4 served as the LLM judge.The response-generation models were limited to two to reduce human-annotation costs.
  • Survey design: Each evaluation used 25 instructions, two candidate responses, one overall preference question, and two aspect questions.Participants also explained why they preferred one response, with response order randomized.
  • LLM judging: The LLM judge used a modified AlpacaEval framework with the same prompts, candidates, ranking questions, explanations, and randomized output order as the SME evaluations.
  • Analysis: Agreement was measured as the percentage of shared LLM-SME preferences, while explanations were analyzed thematically using open coding.

5 RESULTS

Overall SME–LLM agreement was limited but improved modestly when the LLM used an expert persona, while SMEs showed higher agreement with one another.

  • 60% agreement occurred between SMEs and the LLM judge in mental health, compared with 64% in dietetics.
  • 72% SME–SME agreement in mental health and 75% in dietetics established a higher expert-agreement baseline.
  • 4% improvement in SME–LLM agreement occurred in both domains when the expert persona was used for general preference questions.Agreement levels nevertheless remained low.
  • General Preference questions asked participants to select which output was better overall, whereas other categories assessed specific aspects.

5.2 SME vs. LLM-as-a-Judge Agreement on Aspect Questions

SME–LLM agreement varied across aspect questions and domains. Expert personas sometimes improved alignment, but not consistently across aspects.

  • Domain and task complexity affected LLM–SME agreement differently, so expert personas did not consistently improve every aspect.
  • Agreement in aspect questions was lower for dietetics SMEs, with only slight improvement from the expert persona.
  • Mental health SMEs generally showed better aspect-question agreement than dietetics SMEs, except for a notable drop in Clarity agreement.
  • Mental health showed greater LLM alignment than dietetics in Accuracy, Education Context, and Personalization categories.
  • The general model sometimes achieved higher agreement than the expert persona in dietetics.

5.3 SME vs LLM Explanations: Qualitative Results

SMEs prioritized accuracy, professional standards, client-appropriate communication, and personalization, while LLM evaluations often overlooked harmful details or interpreted clarity as comprehensiveness.

  • Alignment with Expert Knowledge and Accuracy: SMEs consistently favored responses containing accurate information and adhering to evidence-based professional standards.This included current dietetics practices and evidence-based mental health treatments.
  • Alignment with Expert Knowledge and Accuracy: LLMs often overlooked harmful or inaccurate content and critical omissions that SMEs identified.They instead repeated details focused on following the prompt instructions.
  • Communication and Clarity: SMEs preferred concise, simple, well-organized explanations, whereas LLMs often equated clarity with greater detail or comprehensiveness.This difference could explain misalignment in Clarity questions, especially when excess information might overwhelm clients.
  • Tone and Framing: SMEs valued positive, empathetic, and encouraging framing that remained supportive rather than judgmental.Mental health experts especially preferred validation, hope, and normalized difficult emotions.
  • Relevance to Client Needs and Actionable Information: SMEs favored personalization that reflected individual preferences, cultural factors, health conditions, and emotional states.LLMs sometimes recognized user fit but often lacked the in-depth rationale provided by SMEs.

6 EVALUATING LAY USER ALIGNMENT WITH LARGE LANGUAGE MODELS

Lay users agreed with the general LLM more often than SMEs did, while expert personas had opposite effects for the two groups. This suggests LLM evaluations may align more closely with lay preferences than specialized expert judgments.

  • The findings emphasize continued SME involvement in developing and fine-tuning LLMs for domain-specific tasks.
  • 80% agreement occurred between lay users and the general LLM in both dietetics and mental health.This exceeded the agreement observed between the general LLM and SMEs.
  • A statistically significant difference was found between lay-user and SME agreement rates with the LLMs (p < 0.0001).
  • Expert personas decreased lay-user–LLM agreement to 76% in both domains.
  • Expert personas improved agreement with SMEs but decreased agreement with lay users.

7 DISCUSSION

The discussion finds that LLM-as-a-Judge alignment with SMEs varies across domains, tasks, and evaluation criteria, limiting the reliability of LLM-only evaluation for expert knowledge tasks. It recommends combining scalable LLM evaluation with targeted SME involvement and domain-specific design.

  • Differences in alignment: LLM–SME alignment varies across domains and aspect questions, even when an expert persona is used.The study reports differences across mental health and dietetics and across criteria such as Accuracy, Education Context, Personalization, and Professional Standard.
  • Persona effects: Expert personas did not uniformly improve alignment: they helped on some specialized tasks but failed to improve dietetics Education Context and Personalization judgments.General-purpose models may perform better when tasks require adaptability and user-centered flexibility.
  • Domain-specific criteria: Mental-health expert personas could favor technical language while missing risks such as confusion, self-diagnosis, or harm in overly detailed outputs.SMEs specifically flagged risks that the persona failed to detect in clarity-related evaluations.
  • Explanation differences: SMEs provided more specific, context-rich explanations than LLMs, which often repeated information from the instructions or outputs.SME explanations contributed new information and insights, whereas LLM explanations tended to be generalized.
  • Design implications: Evaluation frameworks should be tailored to each domain because one-size-fits-all approaches can overlook domain-specific weaknesses.The discussion links this recommendation to variation across criteria such as Accuracy and Education Context.

8 FUTURE WORK & LIMITATIONS

The authors identify several directions for future work, including examining broader domains, pre-training data, and tuning with SME feedback. The study is limited by its two-domain scope and by the lack of further tuning on SME judgments and explanations.

  • Future work: Future work should examine how pre-training data content influences judgment alignment, especially where misinformation or conflicting guidelines are prevalent.Dietetics is identified as an example of a domain with conflicting guidance.
  • Limitations: The evaluated models were not further tuned on pairwise comparisons or qualitative explanations provided by SMEs.The authors note that tuning on specialized feedback requires substantial cost and resources.
  • Future work: Future tuning with SME feedback could allow models to learn from expert judgments and refine their outputs.The paper presents this as a direction for future investigation rather than a demonstrated result.
  • Limitations: The study covers only dietetics and mental health because collecting SME annotations is costly.The authors invite evaluation in fields such as qualitative analysis, creativity, UX, and academic research.

9 CONCLUSIONS

The paper compares LLM-as-a-Judge evaluations with SME evaluations for domain-specific tasks requiring expert knowledge. Its conclusions support SME-in-the-loop workflows, domain-aware evaluation design, and targeted use of expert personas to improve alignment.

  • Conclusion: The paper provides empirical evidence about how LLM-as-a-Judge evaluations compare with SME evaluations on expert knowledge tasks.It examines strengths, weaknesses, and factors that distinguish LLM and SME judgments.
  • Conclusion: The authors conclude that evaluation systems should keep SMEs in the loop for domain-specific tasks requiring expert knowledge.The conclusion frames SME involvement as central to the proposed evaluation design.
  • Conclusion: The proposed implications include SME-in-the-loop evaluation, preference tuning or RLHF, domain-specific frameworks, and expert personas.These mechanisms are presented as ways to improve judgment alignment between human experts and LLMs.

A APPENDIX

The appendix includes an adapted AlpacaEval prompt template and profiles of lay-user participants in dietetics and psychology.

  • Table 4 presents an AlpacaEval prompt template adapted with personification and individual aspect questions.
  • Table 5 profiles lay-user diet and psychology participants using age, sex, ethnicity, education, and related-services client status.
Loading 2410.20266v1…