Source-linked AI summary

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak

arXiv:2609.11067v1cs.CLcs.LG

TL;DR

This paper examines how realistic surface noise affects LLM judges used to measure social bias in text. Across five noise conditions and four judges, noise more often turns neutral judgments into biased ones than the reverse, systematically overestimating measured bias, while robustness reduces the distortion without reversing it.

  • Problem

    Little is known about how typos, informal spelling, and collapsed punctuation affect LLM-based social bias judgments and their implications for bias measurement.

  • Method

    The study applies five rule-based noise conditions at five intensities to 3,822 stereotype-related responses and compares noisy-text judgments with judgments on original text.

  • Results

    Across four judges and twenty-five conditions, neutral-to-biased flips exceed biased-to-neutral flips in every condition, reaching a 120× margin in fragile judges.

  • Takeaways & Limitations

    Bias measured on noisy text is systematically overestimated, especially for fragile judges and categories central to fairness, so judge robustness should be verified.

  • Takeaways & Limitations

    The study does not verify whether a newly selected option is stereotype-aligned, so its fabrication metric captures directional commitment rather than absolute correctness or stereotype alignment.

Abstract

from arXiv · show

Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.

1 Introduction

This paper examines how realistic surface noise affects LLM-as-a-judge measurements of social bias. Across four judges and five noise conditions, noise systematically fabricates bias, especially in fairness-critical categories, while robustness reduces but does not reverse the distortion.

  • Study design: The study applies five surface-noise conditions at multiple intensities to 3,822 stereotype-related responses and compares clean versus noisy decisions across four LLM judges.The conditions include typos, informal spelling, punctuation collapse, a colloquial variant, and a combined condition layering the first three.
  • Main findings: Neutral-to-biased flips exceed biased-to-neutral flips by up to 120× at realistic noise levels, concentrating in religion, disability, socioeconomic status, and sexual orientation.The authors interpret this asymmetry as fabricated bias because surface noise adds no information and preferentially shifts neutral judgments toward directional commitments.
  • Main findings: Fabrication concentrates in the most sensitive bias categories, and this concentration remains stable across noise types.
  • Robustness: Judge robustness varies by roughly an order of magnitude, and increasing robustness attenuates the asymmetry toward parity rather than reversing it.
  • Noise interactions: Combining noise types amplifies fabrication by up to 2.5× over any single type, so single-type robustness checks understate deployment risk.
  • Resource: FABLE2 releases 3,822 binary-choice bias items spanning seven categories with seed-controlled noise injection for reproducible judge stress tests.

2 Related Work

Prior work shows that LLM judges and other language models are sensitive to surface form, prompt wording, and noise. This paper extends that literature by examining how noisy text affects a judge’s assessment of bias in a third party’s response.

  • LLM-as-a-judge: Earlier studies document systematic LLM-judge biases, including position bias, prompt sensitivity, and inconsistent evaluation even on clean inputs.
  • Bias evaluation: Social-bias benchmarks commonly use templated comparisons or ambiguous contexts, while CLEAR-Bias provides the stereotype-probing source material adapted in this study.
  • Surface form: Surface features can alter social judgments: classifiers have flagged African American English as offensive up to twice as often as comparable standard-English text.
  • Noisy text: Existing robustness studies show that typos and punctuation corruption degrade language-model reasoning, including at relatively low noise rates.

3 Experimental Setup

The experiment evaluates four LLM judges on 3,822 labeled stereotype-related responses after applying five surface-noise conditions at multiple intensities. It uses self-baseline flip rates to isolate changes between clean and noisy judgments while controlling for judges’ differing clean tendencies.

  • 3.1 Data: 3,822 samples cover seven bias categories, with 546 samples per category, and each response is labeled as supporting Option A, Option B, or NONE.FABLE recasts CLEAR-Bias prompts into binary-choice items using 91 name pairs for each of 42 questions.
  • 3.1 Data: FABLE uses model-generated responses because they provide controlled free-form text with dependable A/B/NONE labels for varying surface noise in isolation.The authors note that this noise profile differs from authentic user writing.
  • 3.2 Noise Injection: Five perturbations model typos, informal spelling, punctuation collapse, colloquial markers, and a sequential combination of the first three across intensities p ∈ {0.1, 0.3, 0.5, 0.7, 1.0}.At p=0.3, realized token changes range from 16% for punctuation to 47% for the combined condition.
  • 3.2 Noise Injection: The perturbations preserve propositional content by changing orthography and punctuation without adding, removing, or reordering content words.The original stance labels are therefore carried over, although this assumption is not independently verified for every perturbed instance.
  • 3.3 Judges: Four LLMs from different model families judge each clean or noisy response by generating a single A, B, or None label.The judges are Llama-3.1-8B-Instruct, Qwen3-8B, Gemma-4-12B-it, and GPT-5.4.
  • 3.4 Metric: Self-Baseline Flip Rate: Self-baseline flip rate measures the fraction of samples whose noisy-input decision differs from that judge’s own clean-input decision.This avoids entangling noise-induced changes with judges’ differing clean tendencies, while parsing failures count as judge failures.

4 Results

Across judges and noise conditions, surface noise destabilizes bias judgments, with combined noise generally most disruptive and fragile judges especially prone to fabricating bias. Fabrication exceeds erasure most strongly at realistic noise levels, remains concentrated in sensitive categories, and attenuates toward parity as judges become more robust.

  • 4.1 Judges Are Unstable under Surface Noise: 10.7% versus 1.2% was the typo-noise flip-rate range at p=0.3, revealing roughly a 9× robustness gap from Llama to Gemma.Combined noise was the most disruptive condition for every judge, while colloquial and informal noise produced nearly identical rates.
  • 4.2 Noise Fabricates Bias Rather Than Erasing It: 120× more neutral responses became biased than biased responses became neutral for Llama under typo noise at p=0.3, showing strong fabrication asymmetry.Llama produced 240 fabrications versus 2 erasures; under combined noise it produced 288 fabrications versus at most 4 erasures.
  • 4.2 Noise Fabricates Bias Rather Than Erasing It: 120× at p=0.3 versus 50× at p=1.0 shows that Llama’s fabrication asymmetry was purest at realistic, mild noise.At heavier corruption, additional erasures diluted the ratio by scrambling directional judgments.
  • 4.2 Noise Fabricates Bias Rather Than Erasing It: 120×, 5.6×, 3.4×, and 1.3× were the typo-noise fabrication-to-erasure ratios for Llama, Qwen3, Gemma, and GPT at p=0.3.Across 100 judge–condition pairs, erasure exceeded fabrication only eight times, never significantly; robustness removed rather than reversed the distortion.
  • 4.5 The Asymmetry Survives Template-Level Aggregation: 41 of 42 templates favored fabrication over erasure for Llama under combined noise at p=0.3, confirming the asymmetry beyond item-level independence assumptions.Qwen3 showed the same pattern in 39 of 42 templates, while Gemma’s nonsignificant result reflected few flips rather than reversal.

5 Discussion

Noise can push judges from withholding judgment toward premature directional commitments, with the strongest distortion affecting fairness-relevant categories. Robustness attenuates this fabrication toward parity, while combined realistic noise should be part of pipeline stress tests.

  • Category effects: Religion and disability show the highest fabricated bias and mean flip rates under combined realistic noise, while gender shows the lowest.The category pattern is reported for combined noise at p=0.3.
  • Mechanism: Noise blurs contextual cues, encouraging judges to commit to one side from surviving fragments rather than maintain a neutral judgment.The proposed account contrasts neutral decisions, which integrate the whole response, with directional decisions that can rely on partial cues.
  • Robustness: As judge robustness increases, fabrication asymmetry attenuates toward parity rather than reversing.The distortion is strongest in fragile judges and does not become an opposite-direction effect in robust judges.
  • Mechanism: Combined noise removes lexical identity, token-form, and clause-boundary cues simultaneously, amplifying fabrication by leaving judges with sparser evidence.The three perturbations target distinct channels used to recover context.
  • Practical implications: Bias-measurement pipelines should stress-test judges with combined, realistic mild noise and report sensitivity to input quality.The recommendation targets both the noise configuration and the intensity regime where distortion can be especially pure.
  • Scope: Whether fabrication persists on authentic variety text remains open because the controlled Colloquial condition holds grammar fixed.The authors identify authentic variety text as a next question with direct fairness stakes for its speakers.

6 Conclusion

The study tests how surface noise affects LLM-as-a-judge bias measurement across four judges and twenty-five conditions. Noise systematically fabricates bias more often than it erases it, with the strongest distortion in fragile judges and attenuation toward parity as robustness increases.

  • Conclusion: Across four judges and twenty-five conditions, noise fabricates bias rather than erasing it, systematically overestimating bias in noisy text.Fabrication exceeds erasure in every condition, and the overestimation is greatest among fragile judges and fairness-relevant categories.
  • Conclusion: A neutral response is misjudged as biased by up to a 120× margin at mild, realistic noise levels.The imbalance is reported as purest at realistic mild intensity, where erasure is scarcest.
  • Conclusion: As judges become more robust, the asymmetry attenuates toward parity rather than reversing.Robustness removes the distortion without inverting its direction.
  • Implications: Judge robustness should be verified before noisy-text bias measurements are trusted, while authentic user noise remains a promising next test.The conclusion links deployment confidence to robustness checks and identifies authentic user noise as future work.

Limitations

The study’s limitations concern what its transition metric establishes, whether perturbed passages preserve meaning, how judges interpret names, dependence among repeated templates, and the realism and scope of its data and noise.

  • Metric interpretation: The transition metric establishes directional judgment instability, not whether a newly selected option is stereotype-aligned or counter-stereotypical.Stereotype alignment is left to future work and depends on the demographic information carried by the options.
  • Meaning preservation: Meaning preservation was built into the perturbation rules but was not verified through a human study of the noisy passages.The authors note that content words were not added, removed, or reordered, while acknowledging that construction is not verification.
  • Name comprehension: Judges may match corrupted names by orthographic similarity rather than recognize their demographic associations, a distinction requiring targeted controls.Proposed controls replace names with neutral labels, reverse assignments, and directly probe demographic recognition.
  • Repeated templates: The 3,822 items derive from 42 questions and 91 name pairs, so item-level observations are not independent.A template-level sign test preserves the asymmetry, but fuller random-effects modeling would better account for variance.
  • Synthetic noise: Rule-based noise provides controllable intensity but does not fully capture naturally occurring user-generated noise or calibrate its relative frequencies against a reference corpus.The controlled Colloquial condition also holds grammar fixed, unlike genuine non-standard varieties.
  • Scope: All judges evaluate Llama-generated responses selected for balanced labels, so the magnitude may differ for other text sources.The direction is reported as consistent across all four judges, but the source selection may favor unusually ambiguous responses.

B.1 All noise types and intensities

Across all five noise types and five intensities, fabrication significantly exceeds erasure in every condition. Combined noise is the most disruptive condition and produces the largest fabrication counts at every intensity.

  • Overall comparison: Fabrication exceeds erasure significantly in all twenty-five noise-type–intensity conditions.The comparison uses a two-sided binomial test against parity, with all p < 0.001.
  • Overall comparison: Even the mildest condition, punctuation noise at intensity 0.1, fabricates bias 2.5× more often than it erases it.This shows the asymmetry is not confined to maximal corruption or one particular noise type.
  • Combined noise: Combined noise produces the largest fabrication counts at every intensity, reaching 1,210 fabrications against 137 erasures at full intensity.At full intensity, the combined condition produces roughly twice the fabrication count of the typo condition.
  • Intensity: Absolute fabrication generally rises with intensity, except for a slight dip in the Colloquial condition at p=0.7.The intensity trend is reported across all noise types.

B.2 Category ranking is stable across noise types

The category ranking of fabricated bias remains largely stable across typo and combined noise, with only limited reshuffling among the leading categories.

  • B.2 Category ranking is stable across noise types: Only the leading-group order changes: socioeconomic status rises from fourth under typo noise to second under combined noise.Religion and sexual orientation each move down one position.
  • B.2 Category ranking is stable across noise types: Stable composition of the leading group and tail suggests the concentration reflects category properties rather than a specific surface corruption.This interpretation is based on the ranking surviving a change from typo to combined noise.
  • B.2 Category ranking is stable across noise types: The same four axes lead and the same three categories trail under typo and combined noise, despite shifts within the leading group.Disability remains first; religion, sexual orientation, and socioeconomic status occupy the other leading positions, while age, ethnicity, and gender retain the bottom three in identical order.

B.3 Fabrication conditional on eligible cases

Fabrication rates are also evaluated after normalising by the number of clean neutral decisions eligible to become biased, with religion leading under this conditional measure.

  • B.3 Fabrication conditional on eligible cases: Table 10 normalises fabrication by the number of clean NONE decisions available in each category.
  • B.3 Fabrication conditional on eligible cases: Religion has the highest conditional fabrication rate under combined noise at realistic intensity p=0.3.The rate is calculated over the pool of clean NONE decisions available to flip and is summed across four judges.
  • B.3 Fabrication conditional on eligible cases: Religion leads under both conditional normalisation and raw fabrication counts despite having one of the smaller eligible pools.This shows that its leading position is not explained simply by having the largest pool of clean neutral decisions.

B.4 Full transition matrix

The full transition analysis shows that directional reversals are a substantial source of instability, while the fragile judge’s reversals and fabrications amplify its existing directional preference.

  • B.4 Full transition matrix: Directional reversals comprise 39% of Llama’s flips and 16–19% of the other judges’ flips under combined noise.These reversals keep responses biased in both conditions but change which person is implicated, so they are excluded from fabrication–erasure asymmetry.
  • B.4 Full transition matrix: Llama moves A→B 135 times versus 48 in the reverse direction, with 23.4% of clean-A decisions flipping to B compared with 4.2% of clean-B decisions flipping to A.The imbalance persists after normalising by the differing clean decision pools.
  • B.4 Full transition matrix: Llama’s fabrications are likewise lopsided, with 258 transitions toward B versus 30 toward A.Both asymmetries match Llama’s clean tendency to favour B, suggesting noise amplifies an existing directional preference rather than scattering decisions randomly.
  • B.4 Full transition matrix: The full matrix partitions flips into fabrication, erasure, and directional reversal, while GPT-5.4 is omitted because its pipeline lacks per-item labels in the required format.GPT-5.4’s 79 flips nevertheless decompose into 28 fabrications, 38 erasures, and 13 reversals.

C Realized Edit Rates

Realized edit rates differ substantially across nominally matched noise conditions, and fabrication depends on the channels corrupted rather than only on the amount of text changed.

  • C Realized Edit Rates: Punct alters 26% of tokens even at p=1.0, versus 76% for typo, helping explain its lowest flip-rate position.Its character-level distance is non-monotone, peaking at p=0.7 and falling at p=1.0 because deleting all punctuation shortens the string.
  • C Realized Edit Rates: Informal and Colloquial are the only conditions that lengthen text, with noisy-to-clean length ratios of 1.02–1.08.Their length increase is consistent with substituting and multiplying tokens rather than corrupting characters in place.
  • C Realized Edit Rates: Combined produces less realized change than its three components summed, 0.47 versus 0.69 at p=0.3, yet fabricates the most bias.This indicates that fabrication is not determined by total realized change alone.
  • C Realized Edit Rates: Typo at p=0.7 alters 58% of tokens and fabricates 528 cases, while combined at p=0.3 alters 47% and fabricates 530.The near-equal fabrication counts despite different realized changes reinforce that edit volume alone does not explain fabrication.
  • C Realized Edit Rates: The reported edit statistics average token changes, character distance, and noisy-to-clean length over 3,822 responses, while nominal p denotes an edit-attempt probability.Equal nominal intensity therefore does not imply equal realized change across conditions.
Loading 2609.11067v1…