Source-linked AI summary

Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V Chawla, Xiangliang Zhang

arXiv:2410.02736v2cs.CLcs.AI

TL;DR

LLM trustworthiness matters because models can make erroneous judgments based on statistical patterns. This paper presents CALM, an automated framework for quantifying 12 biases in LLM-as-a-Judge, and finds that models can perform reliably on specific tasks while broader judge use still needs improvement.

  • Problem

    Ensuring LLM trustworthiness is important because LLMs may make erroneous judgments based on specific statistical patterns.

  • Method

    CALM is an automated framework that evaluates 12 bias types using automated bias injection and metrics including Robustness Rate and Consistency Rate.

  • Results

    Models may reliably judge specific tasks, but significant room for improvement remains in the broader use of LLMs as judges, with bias effects differing across models.

  • Takeaways & Limitations

    CALM can evaluate future LLM-based judge solutions against higher standards of bias mitigation, while application should account for task-specific bias and model differences.

  • Takeaways & Limitations

    Some question sets and bias-related responses may contain NSFW content, requiring cautious and responsible research use.

Abstract

from arXiv · show

LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many domains, potential issues are under-explored, undermining their reliability and the scope of their utility. Therefore, we identify 12 key potential biases and propose a new automated bias quantification framework-CALM-which systematically quantifies and analyzes each type of bias in LLM-as-a-Judge by using automated and principle-guided modification. Our experiments cover multiple popular language models, and the results indicate that while advanced models have achieved commendable overall performance, significant biases persist in certain specific tasks. Empirical results suggest that there remains room for improvement in the reliability of LLM-as-a-Judge. Moreover, we also discuss the explicit and implicit influence of these biases and give some suggestions for the reliable application of LLM-as-a-Judge. Our work highlights the need for stakeholders to address these issues and remind users to exercise caution in LLM-as-a-Judge applications.

2 PROPOSED FRAMEWORK: CALM

CALM assesses LLM-as-a-Judge bias by applying principle-guided perturbations to judged content or instructions and comparing resulting judgments. Its experiments show that vulnerabilities vary across models and tasks, including dataset-dependent, positional, length, self-enhancement, irrelevant-content, authority, emotional, identity, and CoT effects.

  • Framework components: CALM combines 12 bias categories, diverse datasets, and task-specific metrics to evaluate reliability in LLM-as-a-Judge systems.The framework supports pairwise comparison and scoring judgments across evaluation aspects.
  • Automated perturbation: Principle-guided perturbations modify responses or instructions while preserving correctness and meaning, enabling bias detection through judgment consistency.The judge compares original and modified evaluations; differing outputs indicate bias, while matching outputs indicate robustness.
  • Metrics: Robustness and consistency rates quantify whether judgments remain stable after bias injection or across repeated unperturbed evaluations.CALM also introduces an accuracy metric for CoT bias to measure changes in correct judgments after step-by-step reasoning.
  • Main results: Bias effects differ across models and tasks: Claude-3.5 is generally most resilient, yet GPT-4-Turbo can be inconsistent on emotional responses while ChatGPT is more stable.The results caution against assuming that greater model capability guarantees greater reliability for every bias or task.
  • Main results: Position bias intensifies with more answer candidates, with most models falling below 0.5 robustness when evaluating three or four options.The authors suggest selecting models with stronger robustness rates or randomizing answer order.

5 DISCUSSION

The discussion distinguishes explicit from implicit bias and proposes practical measures for detecting and mitigating bias in LLM-as-a-Judge. CALM’s findings support prompt safeguards and automated detection, while broader reliability improvements remain necessary.

  • Explicit biases are openly stated preferences, whereas implicit biases influence judgments without being acknowledged in the reasoning.Authority bias is given as an example of explicit bias.
  • The authors recommend carefully constructed prompts and advanced reasoning strategies to reduce bias interference.Protective phrases can instruct models to disregard identity information, while step-by-step reasoning can guide judgments.
  • Prompt-injection safeguards are recommended to prevent biased information embedded in prompts from influencing judgments.The proposed safeguards target external attempts to introduce bias into the judging process.
  • A prompt-based detection mechanism can identify potential biases in judging templates before evaluation begins.Its effectiveness varies by bias type, but the authors report promise in uncovering a majority of biases.
  • CALM evaluates 12 bias types through automated bias injection and qualification, providing an objective and scalable assessment approach.The conclusion states that models may reliably judge specific tasks, while broader use still has significant room for improvement.

ETHICAL CONSIDERATION

The paper cautions that some question sets and bias-related responses contain NSFW content. Although the data were manually reviewed and curated, applications should follow ethical guidelines and consider societal impacts.

  • Some question sets and bias-related responses may contain NSFW content.The authors state that the data were manually reviewed and curated for research appropriateness.
  • Applications or extensions should be conducted responsibly with consideration for ethical guidelines and potential societal impacts.

A.1 LLM-AS-A-JUDGE

LLM-as-a-Judge uses language models to evaluate responses without requiring reference texts. This approach has shown performance on open-ended questions that highly matches human preference, while fairness remains an active research focus.

  • LLM-as-a-Judge evaluates responses without requiring reference texts.
  • The method has demonstrated performance on open-ended questions that highly matches human preference.
  • Recent research has examined the fairness of LLM-as-a-Judge.

A.2 FAIRNESS IN TRUSTWORTHY LLMS

Fairness is an ethical principle requiring LLMs to avoid biased or discriminatory outcomes and treat users and groups equitably. Because training-data imbalance can produce demographic biases, fairness affects the trustworthiness of LLM-as-a-Judge.

  • Stereotypes and erroneous judgments based on statistical patterns highlight the importance of fairness in evaluating LLMs.
  • Fairness requires LLMs to avoid biased or discriminatory outcomes and treat users and groups equitably.
  • Pre-training data imbalance can create training imbalances and biases against demographic groups.The paper names gender, age, and language as examples of affected attributes.
  • Fairness in LLMs significantly affects the trustworthiness of LLM-as-a-Judge.

A.3 BIASES IN LLM-AS-A-JUDGE APPLICATION

Prior research identifies multiple cognitive biases that can influence LLM-as-a-Judge evaluations, including positional, verbosity, self-enhancement, order, compassion-fade, egocentric, salience, bandwagon-effect, and attentional biases.

  • Prior studies identify position, verbosity, and self-enhancement biases in LLM-as-a-Judge evaluations.
  • Other reported biases include order, compassion-fade, egocentric, salience, bandwagon-effect, and attentional biases.

B DETAILS OF BIAS TYPES

The paper defines twelve bias types affecting LLM-as-a-Judge decisions and specifies how each can alter evaluations across responses, prompts, identities, references, and reasoning content.

  • Position bias favors responses based on their input position, and its evaluation extends to settings with more than two responses.
  • Verbosity bias reflects potential preferences for longer responses, while compassion-fade bias concerns model-name anonymity in judgments.
  • Bandwagon-effect bias concerns majority opinions, distraction bias concerns irrelevant content, and fallacy-oversight bias concerns recognizing logical errors.
  • Authority bias captures influence from authoritative references, while sentiment bias captures preferences for emotional tones.
  • Diversity bias concerns identity markers; CoT bias concerns explicit reasoning steps; self-enhancement bias concerns favoring a model’s own outputs.
  • Refinement-aware bias concerns different scores for original, refined, and history-exposed answers.

C DETAILS OF BIAS EVALUATION

The evaluation systematically perturbs answers or system instructions to test whether judgments change under controlled bias-related modifications across multiple evaluation settings.

  • Position bias is tested by rotationally permuting answer order for evaluations containing two, three, or four answers.
  • Verbosity bias is tested by lengthening worse answers while preserving essential content, then comparing judgments before and after expansion.
  • Self-enhancement bias compares each model’s scores for its own answers with scores for answers generated by other models without disclosed authorship.
  • Compassion-fade bias compares judgments when model identities are disclosed against judgments under anonymized conditions.
  • Bandwagon and distraction tests insert majority-opinion claims or meaningless system statements and measure resulting judgment robustness.
  • Fallacy, authority, sentiment, diversity, CoT, and refinement-aware tests modify logic, references, emotions, identities, reasoning prompts, or answer histories before re-evaluation.

D DETAILED RESULTS

The detailed results report bias-specific robustness, accuracy, score, and error-rate analyses across judge models, with findings varying by bias and model.

  • Robustness rates are reported for position bias in pairwise and multiple-answer comparisons, and for verbosity across answer-length ratios.
  • Self-enhancement results use Z-score-normalized score heat maps and the ErrorRateSE metric for each judge model.
  • Bandwagon results examine varying public-opinion percentages, whose influence differs across models without a statistical pattern.
  • Distraction results report robustness after irrelevant content is added to both high-quality and low-quality answers.
  • Authority results show that quote and book-type fake references strongly influence most models.
  • Most models do not favor emotionally charged expressions, while CoT and refinement-aware analyses report accuracy and ErrorRateRA metrics.

E CASE STUDY

The case-study section enumerates concrete manifestations of bias using Figures 9–12 and analyzes them in detail.

  • Figures 9–12 present case studies of sentiment, refinement-aware, authority, and bandwagon-effect bias.
  • The case studies illustrate actual manifestations of bias in LLM-as-a-Judge evaluations.
  • The section follows the case studies with detailed analysis of these bias manifestations.

F PROMPT TEMPLATE

This section documents the experimental materials and templates used to evaluate bias across models, datasets, comparison formats, and perturbation types.

  • Table 7 reports detailed experiments for each bias type and lists the corresponding metric values.
  • Figure 8 compares model robustness under sentiment and authority bias, with higher robustness indicating greater resistance to bias.
  • The supplementary materials include prompts for pairwise, triadic, quadruple, and chain-of-thought comparisons.
  • Additional templates generate paired responses, length-expanded sentences, books, URLs, and quotes for bias evaluation.
  • Templates also target compassion-fade, bandwagon-effect, authority, sentiment, diversity, distraction, and refinement-aware biases.
Loading 2410.02736v2…