Source-linked AI summary

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang

arXiv:2608.27309v1cs.CLcs.AIcs.CY

TL;DR

Bounded difference-in-differences endpoints in LLM-judge audits can confound differential preference with differential attenuation. The paper derives this mechanism and tests it in a pre-registered audit of a frozen pedagogy judge, finding a null primary endpoint but a manufactured nominal interaction. Its analysis shows that the censoring contribution is measurable from ratings already collected.

  • Problem

    LLM-judge audits use bounded double differences to certify bias, but the endpoint is not identified when each contrast is censored by its own share.

  • Method

    The paper derives the bounded-scale mechanism and applies it to a pre-registered audit that rates paired scaffolding responses under novice and advanced learner profiles.

  • Results

    +0.085 points was the null primary profile effect (95% BCa [−0.167, +0.353], p = 0.684), while zero differential preference reproduced 79 to 85% of the +0.378 interaction.

  • Takeaways & Limitations

    A non-null bounded difference-in-differences result in an LLM-judge audit cannot by itself be interpreted as preference influence without checking differential censoring.

  • Takeaways & Limitations

    The empirical demonstration is narrow: one judge model, one rubric, six acid-mixture word problems, and 55 stimuli clustered in 23 source tutoring runs.

Abstract

from arXiv · show

Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.

1 Introduction

LLM-judge audits often use bounded difference-in-differences to certify bias, but censoring can make a common severity shift appear as differential preference. A pre-registered pedagogy-judge audit demonstrates this failure while finding a null primary endpoint.

  • 1 Introduction: A double difference contrasts candidate responses within items and then differences that contrast across a manipulated presentation or identity attribute.The design is intended to cancel additive artifacts.
  • 1 Introduction: Censoring each contrast by its own share makes differential attenuation observationally confounded with differential preference.Unequal distances from rating bounds allow a common severity shift to generate an interaction.
  • 1 Introduction: +0.085 points was the registered profile effect on scaffolding preference (95% BCa [−0.167, +0.353], p = 0.684).The estimate was measured against a baseline preference of +2.6 to +3.2 points on the five-point scale.
  • 1 Introduction: +0.378 was the nominally significant productive-struggle interaction (p = 0.002), but zero differential preference reproduced 79 to 85% of its magnitude.The reproduction used the observed severity shift and the scale floor alone.
  • 1 Introduction: The paper contributes a bounded-scale identifiability analysis, a pre-registered audit counterexample, and an audit-computable identifiability check.The analysis carries limited-dependent-variable results into LLM-judge audits.

2 Related work

Related work documents widespread LLM-judge biases, sensitivity to stated attributes, and established nonlinear-scale concerns. This paper focuses on the arithmetic of certifying bias with bounded difference-in-differences.

  • 2 Related work: MT-Bench and Chatbot Arena reported high judge–human agreement alongside position, verbosity, and self-enhancement biases.Later audits added order effects, self-preference, and diverging human and model bias profiles.
  • 2 Related work: Audits of stated attributes find shifts in model predictions, annotation variance, and simulated peer-review outcomes across otherwise controlled conditions.These findings motivate manipulating a stated learner profile.
  • 2 Related work: Pedagogical evaluation research operationalizes scaffolding and productive failure in dialogue benchmarks, remediation corpora, and evaluation taxonomies.The audited instrument is the frozen pedagogy judge of Fan et al. [2026].
  • 2 Related work: Censored and ordinal models show that observed scores are nonlinear transforms of latent quantities, making observed interactions distinct from latent interaction effects.Prior work also shows that observed cross differences need not equal treatment effects in nonlinear difference-in-differences.
  • 2 Related work: Pre-registration is advocated for NLP and predictive modeling because data-dependent analysis can invalidate reported p-values.This study adds a case where registration constrained analysis but an interpretive guarantee remained false.

3 Study design and materials

The study manipulates stated learner profiles in frozen rating contexts rather than generating new tutoring sessions. It compares high- and low-scaffolding responses using a frozen pedagogy judge and controlled materials.

  • 3 Study design and materials: Each frozen stimulus pairs a high-scaffolding response that leaves the next reasoning step to the learner with a low-scaffolding response that performs it.Both responses are rated under dialogue-only, novice-profile, and advanced-profile arms.
  • 3 Study design and materials: A stated learner profile is not a minimal ability-label manipulation because each profile also states a help-seeking preference.One rubric sub-score measures that preference directly, so the endpoint is reported as profile influence.
  • 3 Study design and materials: 55 stimuli were sampled from 2,219 judged tutor turns across six acid-mixture word problems and 23 source tutoring runs.Each pair combines a real next tutor turn with a matched corpus or authored reply.
  • 3 Study design and materials: Competence strata were assigned by majority vote across three independent blind annotation passes, with pairwise inter-pass agreement above 93%.The tutor’s learner-state tracker was used only as a sampling prior to avoid circular labels.
  • 3 Study design and materials: The 59-word parallel profiles were scored by an unchanged frozen Claude Opus 4.8 pedagogy judge on all 990 calls.The judge applied the rubric verbatim across four pedagogical sub-scores.

4 Pre-registered analysis

The registered analysis defines scaffolding preference as the high-minus-low response score and the Profile Anchoring Gap as its profile contrast. Source-run averaging and exact inference address clustered stimuli, while the registration explicitly tests interpretive null branches and assumes censoring cannot manufacture effects.

  • 4 Pre-registered analysis: For stimulus s and arm a, scaffolding preference is Δ_s(a) = S_s(RH | a) − S_s(RL | a).S is the per-unit mean of the judge’s overall rating.
  • 4 Pre-registered analysis: PAG_s = Δ_s(Pnov) − Δ_s(Padv) = aL,s − aH,s, where a_j,s is the profile shift for pole j.The decomposition is an identity, not a model.
  • 4 Pre-registered analysis: All registered estimands are averaged within source run before testing, because stimuli from the same run are not independent.The weak and strong strata span 18 and 10 source runs, respectively.
  • 4 Pre-registered analysis: The primary test is a two-sided exact Wilcoxon signed-rank test on source-run means, with BCa bootstrap intervals based on 10,000 resamples.The test is exact under symmetry about zero and targets a pseudomedian rather than the mean.
  • 4 Pre-registered analysis: A null primary is informative only if the judge demonstrably used the profile, while near-boundary scores can make one-direction effects unobservable.The registration therefore treated a joint null of the primary and pure profile effect as uninformative.
  • 4 Pre-registered analysis: The registration’s claim that censoring cannot manufacture an effect is false for the bounded difference-in-differences endpoint.The paper’s later counterexample limits what a non-null endpoint could license.

5 A bounded difference in differences is not identified

On a bounded rating scale, a double-difference endpoint is not identified because unequal censoring can make a common severity shift appear as differential preference. The paper derives this mechanism and proposes diagnosing its contribution from the audit’s own ratings.

  • A common severity shift can produce a nonzero endpoint when the two candidate-response poles are attenuated unequally.With zero latent differential preference, the observed endpoint equals the difference in attenuation between the poles.
  • PAG = PAG∗ + κHδH − κLδL, so recovering latent preference requires pole-specific censoring shares that the ratings do not supply.The latent endpoint is PAG∗ = δL − δH; the observed rating scale alone does not identify the κj.
  • The artifact is largest when one pole is pinned at a bound and the other remains free, reducing the double difference to a one-pole contrast.At κL = 1 and κH = 0, the endpoint has magnitude |δ| and becomes PAG ≡ −aH.
  • Sharper candidate contrasts generally increase unequal censoring because improved stimuli place the poles near opposite bounds.This design feature drives κH and κL apart rather than eliminating the problem.
  • The pre-registration’s claim that censoring cannot manufacture an effect is false for this endpoint, although the registered primary result remains uncorrupted because it is null.The error changes what a non-null secondary result would have licensed, not the headline primary finding.
  • An audit can estimate the artifact by identifying bound-pinned poles and transporting the free pole’s observed shift onto the pinned pole under zero differential preference.The reproduced magnitude measures what the bounded double difference can generate without differential preference; the construction’s residual is only a prediction error.

6 Results

The stated learner profile barely changes scaffolding preference, but a productive-struggle interaction is largely reproducible from a common severity shift interacting with rating-scale censoring. The audit therefore reports a null primary endpoint alongside evidence that the nominal interaction is not identified as preference.

  • 6.1 The registered primary endpoint is null: +0.085 scale points is the weak-stratum profile effect on scaffolding preference, with 95% BCa [−0.167, +0.353] and p = 0.684.The strong stratum likewise shows +0.106 with p = 0.914.
  • 6.2 The profile influences absolute scores: The profile lowers both poles together, by 0.153 points for low scaffolding and 0.238 points for high scaffolding on weak stimuli.This common-direction movement is consistent with a severity component under unequal attenuation.
  • 6.3 The one significant interaction is not identified: +0.378 is the nominally significant productive-struggle interaction, compared with +0.309, +0.134, and +0.130 for the other three fields.The high pole accounts for 72–96% of total pole movement across fields; for productive struggle, aH = −0.395 and aL = −0.017.
  • 6.3 The one significant interaction is not identified: +0.472 occurs when the productive-struggle low pole is floored, versus +0.149 when it is free.On 17 of 30 weak stimuli, the low pole is exactly 1.000 in all three arms; the resulting gap reduces to the negative high-pole shift.
  • 6.3 The one significant interaction is not identified: +0.321 reproduces 85% of the observed +0.378 productive-struggle gap under zero differential preference, or 79% after integer rounding.The construction transports each observed high-pole shift to the low pole, clips ratings to [1, 5], and re-averages; its residual bounds nothing.
  • 6.4 Censoring is not conservative for this endpoint: Censoring is not necessarily conservative: among weak stimuli with a high pole pinned at 5.000, 1 of 16 fell, versus 8 of 14 when unpinned.The primary mixes −0.205 on 18 pinned stimuli with +0.352 on 12 unpinned stimuli, so the direction of de-censoring is not determined.

7 What the study can and cannot support

The study supports a null profile effect only at its tested resolution, while sparse ratings, multiplicity, and design imbalances constrain interpretation and generalization.

  • Resolution and multiplicity: The test had coarse resolution: five optimally chosen single-rating changes would move the primary endpoint past zero, while six would place it exactly there.The endpoint lies on a sparse lattice because each per-unit score averages three integer ratings.
  • Resolution and multiplicity: Exact tests condition on which clusters moved and discard movement magnitude, so some reported p-values reach their smallest admissible value without rejecting.In four cases, the attainable p-value still exceeds 0.05.
  • Resolution and multiplicity: The 21 pre-registered hypothesis tests and correlated rubric fields make significance and per-field magnitude insufficient to identify the strongest effect.Field exposure to scale bounds differs substantially despite high agreement between some fields.
  • Limitations: The identifiability demonstration is narrow: one judge model, one rubric, six acid-mixture problems, and 55 stimuli clustered in 23 tutoring runs.The response poles also differ in punctuation, boxed answers, tutor-model source, and length.
  • Limitations: The competence strata contain systematic composition differences, including tutor-family confounding and a concentration of single-message contexts in the strong stratum.The all-corpus re-estimate removes the authored-text imbalance but not the single-model imbalance, and remains underpowered.

8 Sharper stimuli make the endpoint less identified

Sharper stimuli can make a bounded-scale difference-in-differences endpoint less identified: pole separation may pin responses near bounds, turning shared severity shifts into apparent interactions.

  • Sharper stimuli make the endpoint less identified: A profile did not detectably shift overall scaffolding preference, but it moved absolute ratings in the same direction on both response poles.The primary null reflects the tested resolution rather than evidence of absence.
  • Sharper stimuli make the endpoint less identified: As stimuli improve, pole separation can convert the within-item difference-in-differences into a one-pole severity contrast because responses approach scale bounds.The same pole separation that supports a manipulation check can increase unequal attenuation.
  • Sharper stimuli make the endpoint less identified: The nominal per-field anchoring effect is mostly reproduced by a model with zero differential preference, linking the apparent effect to attenuation under the bounded scale.The passage frames this as the transferable lesson from the audit.

A Judge instrument and request parameters

The frozen instrument asks the judge to rate only the final tutor turn in context, using four pedagogical dimensions plus an overall score, under profile-arm prompt variants.

  • Instrument: The judge rates pedagogical quality of the final tutor turn in the displayed dialogue, without judging mathematical correctness or tutor identity.Ratings must rely only on the shown dialogue and the move’s pedagogical quality.
  • Instrument: Each rubric dimension uses an integer 1–5 scale, covering scaffolding, productive struggle, assistance calibration, elicitation, and holistic overall quality.The dimensions assess contingent support, preserved student reasoning, amount of help, elicited next reasoning, and overall pedagogy.
  • Instrument: The prompt explicitly avoids rewarding disclosure or withholding categorically, instead requiring judgments calibrated to the student’s state and contribution.A targeted hint may be good pedagogy, whereas generic encouragement that ignores confusion may be poor pedagogy.
  • Request parameters: The three profile arms differ only by inserting background information about the student into the otherwise frozen tutoring-rating prompt.The no-profile arm omits the profile block; novice and advanced texts occupy the profile slot.
  • Request parameters: The design contains 55 × 3 × 2 × 3 = 990 rated calls, with provider defaults, no specified temperature, and a 512-token output cap.Measured token-count differences do not attribute a profile effect to prompt length.
  • Request parameters: The provider returned the same Claude Opus 4.8 identifier on all calls, but no dated snapshot is recoverable and the alias could be repointed.The identifier was constant by construction, not by observation of a dated model version.

B Materials and construction detail

The materials pair high- and low-scaffolding responses across competence strata, but authored-text, tutor-source, and behavioral-context imbalances remain relevant to interpretation.

  • Construction detail: Single-message contexts without tutor turns are concentrated in the strong stratum, leaving systematically thinner behavioral evidence there.There are 16 of 25 strong contexts versus 7 of 30 weak contexts without a tutor turn.
  • Construction detail: Blind competence labeling overturned the sampling prior toward weak seven times and toward strong never, so the strong stratum inherits any prior bias.The passage also notes that the rubric revision was monotone and occurred before judge calls.
  • Registered endpoint: The registered endpoint was +0.085 scale points with 95% BCa [−0.167, +0.353], while the test reached 80% power only beyond approximately 0.40–0.42.The weak stratum used 18 clusters over 30 stimuli; the strong stratum used 10 over 25.

C Registered secondaries and exploratory contrasts

The registered authoring-robustness re-estimate is small and statistically non-significant, while exploratory per-pole contrasts show asymmetric profile effects. The analysis also distinguishes registered status at the estimand level.

  • Registered authoring-robustness re-estimate: +0.0625 with p = 0.750 is the registered authoring-robustness re-estimate on 27 all-corpus pairs.It uses 8 weak-stratum clusters, of which 4 are nonzero.
  • Registered authoring-robustness re-estimate: The robustness estimate is not directly comparable to the primary because its exact test has a much larger attainable p-value floor.The weak-stratum floor is 2/24 = 0.125, versus 2/214 = 1.22 × 10−4 for the primary’s 14 nonzero clusters.
  • Exploratory per-pole contrasts: −0.228 high and −0.281 low are the exploratory Padv − D contrasts on the weak stratum, with p = 0.0469 and p = 0.00195 respectively.The low-pole result reaches its own exact-test floor at 10 of 10 concordant clusters.
  • Exploratory per-pole contrasts: +0.009 high and −0.128 low are the exploratory Pnov − D contrasts, with p = 0.842 and p = 0.0938 respectively.These contrasts show different directions and magnitudes across poles.
  • Specification status: Table 3 assigns specification status to each estimand rather than to an entire analysis section.Registered quantities may appear within exploratory analyses, including per-field decompositions.

D Specification status and additional inferential detail

The paper reports detailed inferential behavior for lattice-supported endpoints and signed-rank power simulations. These results show that interval coverage, test size, and detectable shifts depend on the endpoint and null being evaluated.

  • Lattice-supported endpoints: The per-stimulus PAG endpoint takes eight distinct weak-stratum values, all multiples of 1/3, because each score averages three integer ratings.The source-run means seen by the test lie on a different sparse rational lattice.
  • Inferential detail: Four rows show disagreement between BCa intervals and rank-test decisions because the interval estimates a mean while the test locates a pseudomedian.The interval can exclude zero even when the rank test does not reject.
  • Power simulations: 0.709 is the registered test’s simulated power at the interval upper limit of +0.353, rising above 0.80 between +0.40 and +0.42.The power calculation uses 6,000 nonparametric bootstrap replicates per point.
  • Power simulations: 0.080 under no shift is not the test’s size, because mean-centering an asymmetric distribution violates the signed-rank symmetry null.Under sign flips satisfying the null, the simulated size is 0.049.
Loading 2608.27309v1…