Source-linked AI summary
On the Role of Citations in Preference Data
Yu Hou, Hal Daumé, Rachel Rudinger, William Walden
TL;DR
It is unclear how humans and LLMs evaluate citations when comparing responses, despite citations’ importance for attribution and preference data. The paper uses mixed-effects models on human and four open-source LLM judgments of scientific QA responses, finding that humans prefer diverse but fewer citations while LLM preferences vary by dataset and model. These results motivate more careful treatment of citations in preference-data collection.
Problem
The paper addresses limited understanding of how human and LLM judges weigh citations in pairwise scientific question-answering preferences.
Method
The study applies mixed-effects models to human and four open-source LLM judgments on citation-backed responses from SciArena and ResearchQA.
Results
Humans prefer more diverse but fewer total citations, whereas LLM citation-related preferences differ depending on the dataset and model.
Takeaways & Limitations
Preference-data collection should attend more carefully to citation quality, potentially using detailed instructions, annotation rubrics, external tools, and mixed-effects models.
Takeaways & Limitations
The analysis focuses exclusively on SciArena and ResearchQA, and its mixed-effects results describe correlations rather than causal relationships.
Abstract
from arXiv · showhide
Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs. Yet, it is unclear how humans and LLMs evaluate citations when comparing outputs, a process central to reward modeling and modern LLM post-training. This paper studies the role of citations in the preferences of human judges and four open-source LLMs within the context of scientific question answering, leveraging mixed effects models to investigate the influence of citations on pairwise judgments. Among our key findings are (1) that humans prefer more diverse citations but fewer overall, and (2) that LLMs show some citation-related preferences compared to humans, despite lacking access to the sources, but these preferences depend on the data and specific models. We further discuss the implications of our findings for preference data collection.
1 Introduction
Pairwise preferences are central to LLM post-training and evaluation, but their reasons are often uninterpretable. This paper therefore studies how citation features shape human and LLM preferences for scientific question-answering responses.
- Pairwise preference data supports LLM post-training and evaluation but generally does not reveal why assessors prefer one response.Interpretable preferences are important for understanding bias and developing fair, robust reward models.
- Prior explanation work often uses a propose-and-validate paradigm, but largely neglects long-form responses with citations.
- Citation features increasingly matter in scientific information-seeking as people use AI research assistants to find papers and answer scholarly queries.
- The paper asks how citation-related features influence human preferences and how LLM judges weigh citations in pairwise response assessments.
- Mixed-effects models compare human preferences with LLM judge effects on citation-backed responses from SciArena and ResearchQA.
- Human judges prefer more diverse but fewer overall citations, while four LLM judges show citation and non-citation preferences that differ across models and from humans.
- The findings motivate more careful attention to citations during preference-data collection.
2 General Setup
The study analyzes citation-backed pairwise preferences with mixed-effects models, using selected datasets and LLM judges while controlling for response and presentation differences. Citation diversity, count, age or mismatch, length, and unique content are modeled as predictors of win rate.
- The analysis uses SciArena and ResearchQA, while four open-source LLMs provide additional preference judgments.The judges are Llama3.3-70B, Gemma3-27B, Qwen3-30B, and Skywork-Critic-70B.
- Only datasets with cited responses, pairwise comparisons, and clear human preference labels meet the study’s selection criteria.These requirements constrain the analysis to SciArena and ResearchQA.
- LLM judgments use both response orderings to control for presentation-order bias.
- Mixed-effects logistic regression models binary preferences with fixed effects for predictors and random effects for group-level variation.The models follow the Bradley-Terry framework and use a maximal-then-simplified random-effects design.
- Citation diversity is measured by Shannon entropy, whereas citation count records the total number of citations.Low diversity indicates repeated reliance on a small number of sources.
- Citation age is used only for SciArena, and citation mismatch only for ResearchQA.These predictors reflect dataset-specific citation information and discrepancies.
- Response length and unique content are included as non-citation predictors, with unique content measured using BERTScore precision after removing inline citations.
- Average marginal effects report the change in RA’s win probability, in percentage points, per 1-SD predictor increase.Higher values correspond to more of the relevant citation, response-length, or unique-content property in RA.
3 Case Study: SciArena
SciArena compares human and LLM preferences for citation-related features in pairwise scientific responses using mixed-effects models. Humans favor citation diversity but fewer citations overall, while LLM judges show stronger and more heterogeneous citation and content preferences.
- Setup Details: SciArena contains 12,249 examples from 102 researchers, covering 23 models and 480 unique model pairs after filtering ties, bad responses, and uncited responses.
- Setup Details: The analysis models human and LLM pairwise judgments with dataset-specific mixed-effects structures, including random effects for models, question types, subjects, and continuous predictors.The maximal random-effects structures include random slopes for citation diversity, citation count, citation age, response length, and response content variables.
- Human Preferences: A 1σ increase in citation diversity difference raises the probability that response RA is preferred by humans by 1.8%, whereas a 1σ increase in total citation difference lowers it by 3.5%.Humans also show a slight but nonsignificant preference for older citations.
- Human Preferences: Human citation effects are comparable to response length and semantic content effects, with content uniqueness producing a 4.2 pp preference increase.
- LLM Judge Preferences: LLM judges show stronger preferences for citation diversity, with Δ Win Rate values of 7.5%, 5.2%, 6.2%, and 7.0% for Llama, Gemma, Qwen, and Skywork.Total citation-count effects are mixed: Llama is −2.1 pp and Skywork is −12.8 pp, while the other two effects are nonsignificant.
- LLM Judge Preferences: LLM judges differ in citation and non-citation preferences: citation-age effects are nonsignificant, while content uniqueness correlates more strongly with their judgments than with human preferences.Llama, Gemma, and Qwen also strongly prefer longer responses, unlike Skywork, whose preference is negligible.
4 Case Study: ResearchQA
ResearchQA compares human and LLM preferences for citation-backed scientific responses using mixed-effects models. Humans favor diverse, fewer citations and penalize citation mismatches, while LLM citation preferences vary by model.
- Setup Details: ResearchQA retains 169 human-annotated comparisons spanning eight scientific fields and five general domains, all comparing gpt-4.1-mini with gemini-2.5-flash.
- Setup Details: Mixed-effects models compare human and LLM preferences while modeling response order and citation-related predictors.
- Human Preferences: 11.7% ∆Win Rate favors citation diversity, while -6.5% and -10.1% correspond to citation count and citation mismatch, respectively, in human judgments.
- Human Preferences: Human judges significantly penalize discrepancies between inline citations and reference lists, consistent with concerns about hallucinated citations.
- Human Preferences: Observed human preferences for response length and content differ from SciArena, potentially reflecting differences in annotators, collection setups, and response comparisons.
- LLM Judge Preferences: 7.8 pp and 8.8 pp are significant citation-diversity effects for Llama and Gemma, while Qwen and Skywork show positive but nonsignificant effects.
- LLM Judge Preferences: Llama, Gemma, and Qwen prefer longer responses with ∆Win Rate values of 9.1%, 11.2%, and 5.9%, whereas Skywork does not.
5 Discussion
The discussion argues that underspecified annotation criteria can leave citation-quality judgments vulnerable to heuristics and that LLM judges diverge from humans in data-dependent ways. It recommends clearer rubrics, source access, and more nuanced judge selection.
- Preference protocols often underspecify judging criteria, leaving factors to be learned implicitly by reward models.
- Clear citation-related rubrics may help discourage lazy or heuristic citation-quality assessments by expert annotators.
- LLM judges exhibit substantially different, data-dependent preference patterns from humans, sometimes showing undesirable biases.
- External access to cited documents and domain-specific rubrics are proposed ways to reduce divergences between LLM and human judges.
- LLM judges differ in preferences such as response length, with Skywork lacking the length bias found in three other models.
- Accuracy or correlation with human preferences can obscure individual judge idiosyncrasies, motivating mixed-effects analysis for more principled judge selection.
6 Related Work
Related work covers human-preference evaluation and methods for uncovering latent preferences, while deep-research applications make citation behavior increasingly relevant to meaningful system evaluation.
- Human/LLM Preferences: One line of work evaluates complex tasks using human preferences, including information-seeking through search and coding.
- Human/LLM Preferences: Another line proposes methods to mine or uncover latent preferences from aggregated human pairwise votes.
- Deep Research Applications: Deep-research systems produce long-form, citation-supported responses for challenging tasks, alongside benchmarks such as DeepResearch Bench and ResearchRubrics.
- Deep Research Applications: Understanding citation preferences and underlying human or LLM behavior is presented as essential for more meaningful evaluation of these systems.
7 Conclusion
The conclusion compares citation preferences across human and LLM judges on shared data and identifies model- and data-dependent divergences. It recommends clearer citation instructions, possible external tool access, and mixed-effects analysis for preference-data collection.
- The study compares predictor effects between human and LLM judges and across different LLM judge types using the same data.
- Humans prefer diverse but fewer citations, whereas LLM citation preferences depend on the dataset and model.
- Some LLM judges share identifiable biases, while humans may display behaviors such as penalizing mismatched citations that LLM judges lack.
- Different training processes appear associated with different LLM preferences for features such as response length.
- Detailed citation-quality instructions, annotation rubrics, external tool access, and mixed-effects models are suggested for preference-data collection.
Ethics Statement
The authors report no substantive ethical concerns because the study uses public, anonymized datasets without personally identifiable information and evaluates open-weight models observationally.
- The analysis uses publicly available, anonymized datasets containing no personally identifiable information.
- The study evaluates open-weight models and is observational, aiming to understand citation-related preferences.
Limitations
The study’s findings are bounded by dataset selection, prompt sensitivity, possible nonlinear relationships, and the correlational nature of the mixed-effects results.
- Dataset: The analysis uses only SciArena and ResearchQA, and different datasets may yield different observations about human preferences.The authors believe broad dataset coverage and a widely applicable modeling approach support likely generalization.
- Prompting: Different prompts could produce variations in observed model preferences, although tested prompt variants had highly correlated results.The authors defer substantially different prompts and their effects to future work.
- Linearity Assumption: The linearity assumption may miss more complex relationships or interactions among variables.The authors found no evidence of violation but encourage future work on nonlinear and interaction effects.
- Correlation vs. Causality: The mixed-effects results reflect correlations rather than causal relationships between features and preferences.The authors note that LLM judges may still learn to rely on non-causal features.
A.1 LLM Judge Experiment Setup
The LLM judge experiments use zero-shot prompting with the SciArena setup, evaluate four open-weight models, and control for response presentation order.
- The experiments use the same zero-shot prompt setup as SciArena.
- The evaluated models are Llama3.3-70B, Gemma3-27B, Qwen3-30B, and Skywork-Critic-70B.
- The models are served with vLLM using NVIDIA A100 GPUs and a temperature of 0.8.
- Both response presentation orders are tested to control for potential positional bias in citation preferences.The flipped-order experiment switches the order of responses while retaining the same prompt template.
- The prompt asks judges to assess relevance, accuracy, clarity, and citation use before selecting the better response.
B MEM Details & Model Selection
The paper fits complex mixed-effects models with group-level random effects and continuous random slopes, simplifying structures when necessary for valid fits.
- The modeling strategy starts with a maximal random-effects structure and removes terms to ensure a non-singular fit.The final structure is further simplified based on AIC.
- The complex modeling approach is computationally costly and may over-specify the data-generating process.
- SciArena groups observations by models, query types, query subjects, and model pairs, with 23 models, 6 query types, 25 subjects, and 480 model pairs.
- The SciArena model includes random intercepts for models, model pairs, query types, query subjects, and model-by-query interactions.
- ResearchQA uses a reduced maximal form with a random intercept for query field because it has fewer data groups and 169 observations.