Source-linked AI summary
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
TL;DR
Existing measures cannot disentangle genuine LLM self-preference from response quality, stylistic cues, and evaluator uncertainty. Using narrative selections under blind and labeled evaluation, the study finds that controlled self-preference largely disappears, while attribution labels alone produce bidirectional bias.
Problem
Existing self-preference measures conflate response quality and surface cues, leaving whether the bias persists under control and whether labels alone induce it unresolved.
Method
Ten LLM judges evaluated narrative constraint selections from a fixed pool under blind and labeled conditions, controlling for selection quality and judge severity.
Results
Controlled blind evaluation largely eliminates self-preference, whereas labels alone systematically inflate self-labeled scores and deflate other-labeled scores regardless of source.
Takeaways & Limitations
Authorship attribution is a distinct source of LLM judge bias, and open-ended ground-truth-free tasks can provide controlled instruments for studying it.
Takeaways & Limitations
The findings come from a single creative-selection domain, and explicit attribution prompts may not generalize to other evaluations or real-world deployments.
Abstract
from arXiv · showhide
As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
1 Introduction
The study tests LLM self-preference with a creative constraint-selection task that removes stylistic confounds and controls selection quality and evaluator severity. It finds that apparent self-preference largely disappears under blind evaluation, while self/other labels alone induce bidirectional bias under matched quality.
- Experimental design: Ten LLMs select narrative constraints from a pool of 200 and judge the selections on four rubric dimensions under blind and experimentally manipulated labeled conditions.The labeled condition uses a 2×2 design crossing true versus false labels with claimed self versus other sources.
- Methodological confounds: Controlling selection quality and judge severity makes apparent self-preference largely disappear, indicating that prior measurements may conflate preference with response quality and evaluator uncertainty.Earlier work attributed almost 90% of observed self-preference to evaluator uncertainty, a judge’s limited ability to assess response quality reliably.
- Blind evaluation: Under blind evaluation, self-preference vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original.Most LLMs initially assign higher mean scores to their own selections, but this pattern does not persist after the stated controls.
- Labeled evaluation: Under matched quality, self/other labels alone inflate scores for self-labeled selections and deflate scores for other-labeled selections regardless of the selections’ actual source.This provides direct evidence that authorship attribution is a distinct driver of evaluation bias.
- Contribution: Open-ended, groundtruth-free creative tasks can function as controlled instruments for studying LLM judge biases without model-specific stylistic fingerprints.The selections retain a recoverable model-specific signature while avoiding the stylistic features of generated text.
2 Related Work
Prior work establishes that LLM judges are widely used but vulnerable to systematic biases that threaten evaluation reliability. Research links self-preference to competing mechanisms, while authorship-label effects remain confounded by model reputation.
- LLM-as-a-judge biases: LLMs are increasingly deployed for pairwise comparison and direct scoring, but systematic biases such as position bias can compromise evaluation validity.Position bias favors responses based on order rather than quality.
- Self-preference: Self-preference has been demonstrated across diverse evaluation domains, but its underlying sources remain contested.Proposed explanations include links to self-recognition ability and familiarity with lower-perplexity outputs.
- Authorship-label bias: Authorship labels can shift LLM evaluations, with judges sometimes attending more closely to labels than content and exhibiting label preference hierarchies.Reported hierarchy: Expert > Human > LLM > Unknown.
- Authorship-label bias: Because prior studies used real model names, observed evaluation shifts may reflect model reputation effects rather than self-preference.This ambiguity motivates separating authorship attribution from reputation in bias analyses.
3 Methodology
The methodology replaces text evaluation with structured narrative-constraint selections, using ten LLMs as selectors and judges across blind and labeled experiments. It establishes recoverable model-specific selection signatures and estimates self–other and label effects with quality-matched designs and mixed-effects controls.
- Evaluation targets: Ten LLMs each selected 20 of 200 narrative constraints across 30 runs, producing 300 evaluation targets assessed on a four-dimension rubric.The rubric dimensions were Originality, Dimensionality, Coherence, and Tellability, scored on a 7-point Likert scale.
- Model-specific signatures: 50.7% leave-one-out k-NN accuracy identified source models from selections, exceeding the 10% chance baseline; within-model Jaccard similarity also exceeded chance for all ten models.Accuracy remained 46.7–50.7% across k ∈ {1, 3, 5, 7}.
- Experiment 1: In Experiment 1, each judge evaluated all 300 unlabeled selections three times, yielding 9,000 runs with independently shuffled constraint and rubric-dimension orders.The design includes a raw self–other comparison and a confound-controlled analysis that accounts for judge severity, selector quality, and repeated-selection nonindependence.
- Experiment 2: A mixed-effects model separated displayed-label effects from actual-source effects and their interaction, while baseline deviations decomposed whether labels inflated self scores, deflated other scores, or both.The displayed-label coefficient captures the average self-versus-other score shift, whereas the actual-source coefficient captures differences between self- and other-produced selections.
4 Results
Results show that apparent self-preference under blind evaluation is largely explained by selection quality and low-discrimination judges, whereas matched-quality labels alone produce bidirectional attribution bias. Displayed labels significantly shift evaluations even when actual source does not.
- Experiment 1: Blind evaluation: 8 of 10 judges produced meaningful score variation, while Llama 4 Maverick and Mistral Large 3 lacked sufficient discrimination for self–other comparisons.Both low-discrimination judges are open-weight LLMs.
- Experiment 1: Blind evaluation: 7 of 10 judges rated their own selections higher on average, but the gap tracked selection quality rather than genuine self-preference.Kimi K2.6 had the largest raw self–other gap while ranking among the highest from nearly every judge.
- Experiment 1: Blind evaluation: After excluding two low-discrimination judges, positive self-preference effects disappeared on the average score and three dimensions; only Originality remained negative.The surviving negative Originality effect runs counter to self-preference.
- Experiment 2: Matched quality: 0.29–0.57 self-label increases occurred across all four dimensions within each actual-source block, including Originality, showing that labels shifted evaluations of identical selections.The shift was attributed to displayed labels rather than content.
- Experiment 2: Matched quality: βL was significantly positive across all dimensions and the average score, while βA and βLA were nonsignificant, confirming that matching removed residual quality differences.The label effect remained robust after excluding two low-discrimination judges.
- Experiment 2: Matched quality: 9 of 10 judges responded differentially to displayed labels, with 6 inflating self-labeled scores and deflating other-labeled scores.Claude Opus 4.7 and DeepSeek-V4-Pro showed self-boost only; GPT-5.5 showed other-penalty, while Llama 4 Maverick deflated under both labels.
5 Conclusion
The study uses a stylistically controlled narrative-selection task to test self-preference in LLM judges under blind and labeled evaluation. Self-preference disappears after controlling for quality and judge severity, while disclosed authorship labels systematically shift evaluations regardless of the target’s actual source.
- Study design: A narrative selection task eliminates stylistic features by design, enabling controlled study of LLM judge behavior under blind and labeled evaluation.The study examines whether self-preference persists after surface-level cues and confounders are controlled, and whether authorship labels alone induce bias.
- Blind evaluation: After controlling for output quality and judge severity, apparent self-preference disappears in blind evaluation.Although judges initially assign higher mean scores to their own selections, the effect does not persist under these controls.
- Blind evaluation: On the only significant dimension, Originality, controlled blind evaluation produces self-deprecation rather than self-preference.Judges rate their own selections less favorably on this dimension after control.
- Labeled evaluation: Disclosed authorship labels systematically shift judges’ evaluations regardless of the evaluation target’s actual source.This finding indicates that the label itself can affect judgments independently of the selection’s true authorship.
6 Discussion
The discussion identifies authorship attribution as an independent source of LLM-judge bias: self-labels inflate scores while other-labels deflate them, even without stylistic self-recognition cues. The reversal of Originality effects across experiments and the ambiguity of creative evaluation motivate further study of how judges operationalize attribution.
- Self-labels inflate scores and other-labels deflate them, indicating bidirectional label-induced bias not fully explained by familiarity-based self-preference.The effect arises from labels alone rather than properties of evaluated content.
- Explicit attribution labels can induce systematic score inflation and deflation even when stylistic signals supporting implicit self-recognition are structurally removed.This suggests the self-recognition framework may extend beyond internal cues.
- Because labels established attribution categories without revealing model identities, how judges latently operationalize those categories remains an open question.The experiments demonstrate the distinction without disclosing actual model identities.
- Originality reversed across experiments: Experiment 1 showed significant self-deprecation after controls, whereas Experiment 2 showed strong label-induced inflation.The reversal supports attribution-sensitive evaluation behavior rather than a uniform self-preference effect.
- Authorship labels may have especially pronounced effects in creative evaluation, where ambiguity requires judges to resolve multiple plausible interpretations.Such labels may shape both overall scores and specific dimensions, as illustrated by the Originality reversal.
Limitations · A Model Setup
The study’s limitations concern domain coverage and ecological validity, while its model setup uses a heterogeneous panel of models queried under standardized settings. Generalization beyond creative selection and explicit attribution prompts remains unresolved.
- Limitations: The findings are limited to a single creative-selection domain.Whether similar attribution effects extend to reasoning or instruction-following evaluations remains unclear.
- Limitations: Attribution-effect generalization to reasoning evaluations remains unresolved.The passage specifically identifies reasoning as a setting requiring further study.
- Limitations: Attribution-effect generalization to instruction-following evaluations remains unresolved.Instruction-following is likewise identified as an untested LLM-as-a-judge setting.
- Limitations: The labeled evaluation setup relies on explicit self/other attribution prompts.This design may reduce ecological validity relative to real-world deployments.
- Limitations: Real-world LLM-as-a-judge deployments often make authorship information implicit or unavailable.This difference motivates caution when transferring the study’s labeled-setting findings to deployment contexts.
- A Model Setup: The model set spans nine developers and includes frontier general-purpose and open-weight systems.The models vary in scale and training lineage, providing heterogeneity for probing self-preference across model families.
- A Model Setup: All models were accessed via OpenRouter1 with exact snapshot identifiers provided.This documents the access route and model versions used in the setup.
- A Model Setup: Models were queried with temperature = 1.0 and reasoning_effort = high where supported.Unsupported parameters were left at their defaults.
B Evaluation Rubric · C Evaluation Run Template
The paper evaluates narrative selections with a four-dimension, selection-level rubric scored independently on a 7-point scale. The rubric is a fixed comparative instrument rather than a human-validated measure of absolute quality.
- B Evaluation Rubric: The rubric scores selections on Originality, Dimensionality, Coherence, and Tellability rather than assigning one overall quality score.Each dimension targets a distinct property of narrative design.
- B Evaluation Rubric: Originality and Coherence adapt the Torrance Test of Creative Writing from surface text to configurations of selected narrative components.Originality rewards departures from familiar patterns, while Coherence assesses whether the components can form one story.
- B Evaluation Rubric: Dimensionality measures how narrative components interlock to create depth, with each element shaping the meaning of the others.It captures constitutive interdependence among selected components.
- B Evaluation Rubric: Coherence asks whether components can share one story, whereas Dimensionality asks whether they need one another once they do.The distinction separates narrative compatibility from mutual meaning-making.
- B Evaluation Rubric: Tellability measures whether a selection has a conflict, question, or stake that gives readers a reason to care.A selection may be original, coherent, and dimensional while still lacking a compelling point.
- B Evaluation Rubric: Selections are scored individually on a 7-point scale with described endpoints rather than compared pairwise.Direct scoring provides an absolute number that allows label manipulation to shift a single rating in Experiment 2.
- B Evaluation Rubric: The rubric was not validated against human annotation and therefore serves only as a held-constant comparative frame.Human validation would be required before using it to assign absolute quality scores.
C.1 Experiment 1: Blind Evaluation · C.2 Experiment 2: Labeled Evaluation
The experiments evaluate narrative-constraint selections on four 1–7 dimensions, first blindly and then under manipulated self/other labels. Selection profiles also retain source-model information substantially above chance, enabling labeled evaluation without model names.
- C.1 Experiment 1: Blind Evaluation: The blind-evaluation setup recorded the judge, selector, self-status, repetition, and dimension order for each run.One example lists Kimi K2.6 as judge, GPT-5.5 as selector, and self-status as False.
- C.1 Experiment 1: Blind Evaluation: Experiment 1 had judges rate 20 narrative constraints on originality, coherence, tellability, and dimensionality using 1–7 Likert scores.The dimensions were presented in randomized order.
- C.1 Experiment 1: Blind Evaluation: Each dimension required a score plus 2–3 English sentences explaining the rating, with instructions to use the full 1–7 range.
- C.2 Experiment 2: Labeled Evaluation: Experiment 2 changed only the opening label phrase, allowing a selection from another model to be labeled as the judge’s own under the FL-self condition.The schematic example shows Gemini 3.1 Pro evaluating a Qwen3.6-Plus selection.
- C.2 Experiment 2: Labeled Evaluation: 0.507 observed k-NN source-classification accuracy exceeded the 0.10 chance baseline, with a null mean of 0.095 and p < .001.The null distribution used 5,000 random label shuffles and leave-one-out k-NN with k = 5.
- C.2 Experiment 2: Labeled Evaluation: Leave-one-out k-NN classification used N = 300 selections across 10 classes, with Cohen’s κ of 0.41–0.45 indicating moderate chance-corrected agreement.The primary configuration was k = 5, tested against H0: p = 1/10.
- C.2 Experiment 2: Labeled Evaluation: Llama 4 Maverick and GPT-5.5 had perfect recall but lower precision, whereas Mistral Large 3 had perfect precision but was rarely predicted.These are the two asymmetric patterns highlighted in the per-model report.
E Within-Model Jaccard Similarity · F Self–Other Comparison by Judge Across Rubric Dimensions · G Average Score by Judge and Selector
The analysis measures within-model selection consistency against a random-selection null, then examines raw self–other differences and judge–selector score patterns across rubric dimensions. The score patterns indicate that selection quality and judge severity contribute to observed diagonal self-evaluation gaps.
- E Within-Model Jaccard Similarity: 435 unique selection-run pairs were evaluated for within-model consistency using pairwise Jaccard similarity.The pairs came from 30 selection runs per model.
- E Within-Model Jaccard Similarity: The random null had a mean Jaccard similarity of 0.054 with a 95% CI of [0.051, 0.058].The null was estimated from 10,000 permutations of uniform random selection from 200 constraints.
- E Within-Model Jaccard Similarity: Significance was defined relative to the random null using a one-tailed permutation test, with ∗ denoting p < .0001.The reported threshold applies to comparisons against the random null.
- F Self–Other Comparison by Judge Across Rubric Dimensions: The raw self–other comparison reports ∆ = Self − Other across all four rubric dimensions without controlling confounds.This comparison is organized separately for each judge.
- G Average Score by Judge and Selector: Figure 9 organizes average scores by judge in rows and selector in columns, with diagonal cells representing self-evaluation.Row variation reflects judge severity, whereas column variation reflects selection quality.
- G Average Score by Judge and Selector: Selectors receiving consistently high scores across judges, notably GPT-5.5 and Kimi K2.6, also score highly on the diagonal.This pattern suggests that selection quality partly drives the raw self–other gap rather than genuine self-preference alone.
H Experiment 2 Estimates Excluding Low-Discrimination Judges
After excluding Llama 4 Maverick and Mistral Large 3, the displayed-label effect remained the only significant fixed effect across every Experiment 2 dimension, while actual-source effects were expected to be near zero under successful quality matching.
- Robustness check: The robustness check excluded the two low-discrimination judges, Llama 4 Maverick and Mistral Large 3.The mixed-effects model specification was identical to Table 5.
- Fixed effects: βL, the displayed-label effect, remained the only significant fixed effect on every dimension.The passage reports ∗∗∗p < .001 for significance.
- Model terms: βA measured the actual-source effect and was expected to be approximately zero under successful quality matching, while βLA captured their interaction.The table defines βL as the displayed-label effect, βA as the actual-source effect, and βLA as their interaction.