Source-linked AI summary
Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews
Jiabin Zheng
TL;DR
The paper asks whether human reviewers changed what they reward when lexically elaborate prose became cheaper to produce, a question ordinary score–text trends cannot identify. It uses a fixed-prompt, fixed-model rater as a composition comparator across ICLR submissions, finding that human valuation of non-domain lexical complexity declined while the frozen rater preserved its earlier schedule. The findings support a reviewer-preference shift and warn that historically calibrated LLM judges can silently misalign despite ordinary total-score agreement.
Problem
Whether reviewers changed their valuation of lexical complexity after its production cost collapsed is unknown because score–text trends confound evaluator change with submission composition.
Method
A fixed-prompt, fixed-model rater retrospectively scores 2018–2025 ICLR submissions, using its stable coefficients to separate composition change from human preference drift.
Results
Human valuation of non-domain lexical complexity declined while the frozen rater preserved its earlier reward schedule, with the divergence surviving placebo and specification checks.
Takeaways & Limitations
Reviewers discounted a cue whose production cost collapsed, while historically calibrated LLM judges can silently drift from contemporary preferences despite ordinary total-score agreement.
Takeaways & Limitations
The evidence covers abstracts from one venue, one model family under one prompt, and eight years, with the comparability condition tested rather than guaranteed.
Abstract
from arXiv · showhide
Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.
1 Introduction
The study asks whether human reviewers changed their valuation of lexical complexity after its production cost collapsed, separating evaluator change from shifting submission composition with a frozen rater. It finds divergence between human and frozen-rater trends, with robustness checks and a submitted-text correction supporting the identified shift.
- Motivation: A falling score coefficient cannot distinguish changed evaluator weights from changed text distributions or other differences among papers carrying the cue.The paper frames this as an identification problem rather than a simple trend in text-score associations.
- Identification strategy: The frozen-rater design uses one model family and prompt applied retrospectively to 2018–2025 submissions, so its coefficient movement tracks composition while the human-minus-frozen trend identifies human change.The machine is treated as a fixed ruler, not ground truth, and the comparison is implemented as a feature × year × human three-way difference-in-differences.
- Robustness: Forty random-wordlist placebos, a year permutation, and a pre-period window leave the lexical-complexity divergence intact, while sentence-length variability remains a human-only reward.The study applies false-discovery control and interval exclusion to its claims.
- Data provenance: Rebuilding the corpus with submitted rather than post-revision abstracts makes the decline steeper rather than weaker.Revision is more common for accepted papers and later years, motivating the data-provenance correction.
- Interpretation: Reviewers discounted a cue whose production cost collapsed, whereas the frozen rater preserved its earlier reward schedule and can drift from current preferences despite ordinary total-score agreement.The paper presents this as consistent with models of evaluation under manipulable signals and as a warning for historical-preference-calibrated LLM judges.
2 Related work
Related work documents changing writing–evaluation associations, biases in LLM judges, measurement-invariance methods, and economic models of signal debasement. This paper positions its frozen rater as an instrument for identifying whose evaluation standard moved, not as a superior judge.
- Observed shifts: Prior ICLR work reports that the notion of good writing shifted, but its period correlations lack a fixed comparator and therefore leave the source of movement unidentified.The paper replicates that design while treating identification, rather than trajectory description, as its contribution.
- Machine raters: LLM-review research documents verbosity, position, fairness, cognitive, self-preference, and hidden biases, including overemphasis on stylistic signals.Controlled edits also show that LLM reviewers can be manipulated by surface changes.
- Measurement invariance: The paper uses measurement-invariance vocabulary: non-uniform differential item functioning concerns group differences in slopes and requires invariant anchor items.The frozen judge serves as the anchor for human-review changes across years.
- Signal debasement: Signal-debasement models predict that receivers underweight costly cues when senders can manipulate them, increasingly so as manipulation becomes cheaper or more heterogeneous.The paper treats this economic framework as an interpretation of an identified shift rather than as an assumption generating the result.
3 Method
The method uses submitted-version abstracts and a frozen machine rater to separate evaluator movement from composition drift while estimating cue effects on standardised review scores. It combines within-year controls, a three-way trend comparison, invariant-rater assumptions, and placebo-based validation.
- Corpus and measurement: 32,638 submissions, 124,615 human reviews, and 81,850 frozen machine reviews form the corpus spanning ICLR 2018–2025.The human and machine arms cover the same eight ICLR cycles, with differing machine coverage by year.
- Corpus and measurement: Submitted-version abstracts are used because current released text differs from what reviewers read, and full-text alternatives are outcome-selected.Matching to arXiv or OpenAlex would skew samples toward accepted papers, while the venue API would require 38 days for the corpus.
- Corpus and measurement: Four cues are mutually controlled: non-domain and domain lexical complexity, sentence-length variability, and abstract length.Lexical complexity is split so general scientific vocabulary measures style while domain jargon measures topic; abstract length controls a known confound.
- Composition drift: 32.67 to 43.18: non-domain lexical complexity rises sharply after 2023, while abstract length rises from 161 to 189 words and sentence-length variability falls from 7.85 to 7.60.Domain jargon also rises from 47.39 to 51.34, establishing substantial composition drift across cycles.
- Identification: The estimand is the feature × year × human interaction, a multi-period difference-in-differences contrasting human and frozen-rater trend changes.Effect sizes are reported as score-standard-deviation changes from moving a cue from its median to its 90th percentile, with bootstrap intervals at B = 1,500.
- Identification: The frozen rater tracks composition alone because one model family, prompt, and short generation window keep its scoring rule fixed across years.The design requires the machine rater to remain unrecalibrated, not to agree with or accurately imitate humans.
- Validation: Forty random mid-frequency wordlists and thirty year-label permutations provide placebo null distributions for the identical comparative specification.The design also requires both arms to score the same papers and assumes composition changes affect both arms in the same direction.
4 Results
Human reviewers’ weighting of lexical cues declined while the frozen rater remained comparatively stable, whereas sentence-length variability showed the opposite pattern. Robustness checks support a cue-specific human shift rather than generic shrinkage.
- Lexical-cue trends: The human coefficient on non-domain lexical complexity fell from +0.142 to −0.015, while the frozen rater changed from +0.080 to +0.082.The inference rests on the trend difference rather than any single noisy yearly coefficient.
- Lexical-cue trends: Domain jargon declined in the human arm from +0.085 to −0.013, while the frozen arm remained near a low constant.
- Cue specificity: Sentence-length variability remained positively weighted by humans but stayed near zero for the frozen rater throughout.The human decline was therefore confined to the two lexical cues.
- Inference: The three-way difference-in-differences was −0.0100 for non-domain lexical complexity and −0.0098 for domain jargon, with both clearing the double gate.Sentence-length variability instead produced −0.0005 and failed the gate.
- Robustness: Forty random-wordlist placebos centred on −0.0005, while the real lexical-cue estimates fell outside the placebo distribution.Year-label permutations likewise produced interactions centred near zero.
- Robustness: The effect was concentrated in 2022–2025, with −0.0353 compared with −0.0129 in 2018–2021.
5 Discussion, limitations, and conclusion
The study interprets the human–frozen-rater divergence as reviewer preference change, while cautioning that the frozen rater is an instrument rather than a universal judging standard. Its scope is bounded by specific data, model, venue, and measurement limitations.
- Discussion: Human reviewers reduced the weight on a cue whose production cost collapsed, consistent with manipulable-signal models but not proving deliberate optimization.The timing suggests reviewers moved before the submitted papers changed, while the authors cannot distinguish optimization from an ordinary taste change.
- Discussion: The frozen rater retained an earlier reward schedule, paying +0.0715 for lexical complexity while missing the structural feature humans now price.Its total-score agreement with humans was 0.245, close to human–human agreement, making the misalignment invisible from agreement statistics alone.
- Implications: Historical-preference calibration can preserve obsolete lexical rewards; judges calibrated earlier should place larger weights on non-domain lexical complexity than contemporary reviewers.Judges refreshed on recent reviews are expected to show a smaller gap, measurable with the same design.
- Limitations: The evidence is limited to abstracts from one venue over eight years, lexical and structural proxies, and one model family under one prompt.Human ratings are noisy, and version effects remain for 2024 abstracts and machine scoring of current PDFs.
- Conclusion: A fixed automated rater can diagnose human preference change, but it should not itself be mistaken for the standard being measured.The rater must remain unchanged throughout the retrospective study; live or silently rerouted endpoints do not satisfy this requirement.
6 Reproducibility
The study makes its data, code, and frozen-rater configuration publicly checkable through released inputs, snapshots, scripts, and configuration fields.
- Data availability: Human reviews, decisions, abstract histories, machine reviews, and pre-review snapshots are publicly identified as the study’s data inputs.The sources include OpenReview, the Gen-Review corpus, and an ICLR dataset repository.
- Code availability: The released pipeline includes corpus assembly, abstract-version rebuilding, feature construction, estimation, placebo and permutation draws, diagnostics, and rendering scripts.The release is intended to reproduce every table and figure from a single machine-readable pipeline.
- Rater configuration: The frozen arm is specified by its model family, generation window, prompt, and decoding settings, reproduced from the Gen-Review release.These fields are the basis for verifying that the rater remained fixed.
A All pre-specified tests
The pre-specified testing framework reports which headline inferences survive both false-discovery control and interval exclusion, while separating estimands with distinct scales and units.
- All pre-specified tests: 23 pre-specified tests are marked as passing only when BH-FDR q < 0.10 and the 90% interval excludes zero.The subscore families do not dissociate on the two lexical cues after author count is controlled.
- Headline inference: Figure 7 displays nine headline rows with 90% intervals and BH-adjusted q values, using separate scales and units for each estimand.Filled markers pass both gates; open markers do not, with colour and shape distinguishing raters.
B Withdrawn findings
The paper reports seven findings withdrawn after adversarial re-testing rather than presenting them as retained evidence.
- Withdrawn findings: Seven findings were withdrawn after adversarial re-testing, with each row documenting the original claim, overturning diagnostic, and outcome.The diagnostic scripts are released.
C Abstract versions
Table 10 identifies the text used to compute the cues and documents the rebuilt abstract layer and version audit that motivated it.
- Table 10 identifies the text on which the cues were computed.
- The table documents a rebuilt abstract layer.
- The table records the version audit that motivated the rebuilt abstract layer.
D Explained variance
The four-cue model’s explained variance is reported by ICLR cycle and rater. Figure 8 presents yearly R2 descriptively, with recent cycles shaded rather than fitting trends or testing identification.
- Explained variance is tabulated by year and rater for the four-cue model.
- Figure 8 reports observed yearly R2 for the four-cue model across eight observations.
- The 2023–2025 cycles are shaded in the figure.
- The figure’s lines join observations without fitting trends and provide a descriptive diagnostic, not an identification test.