Source-linked AI summary

Do Large Language Models Favour Any Research Topics?

Mike Thelwall

arXiv:2609.00323v1cs.DL

TL;DR

LLM-based article scoring may support research-quality evaluation, but evidence about topic and input biases remains limited. This study compares two LLMs and two input formats across health and life sciences articles, finding systematic differences in favored topics and scores that support treating LLM bias as a substantive evaluation concern.

  • Problem

    Statistically strong evidence was lacking on whether LLMs systematically favor particular article types when evaluating research quality.

  • Method

    The study compared GPT-OSS-120B and Gemma 3 27B scores for health and life sciences articles using title/abstract and full-text inputs, analyzing word-frequency differences.

  • Results

    The analyses found systematic topic differences between LLMs and systematic score differences between full-text and title/abstract inputs.

  • Takeaways & Limitations

    LLM biases should be considered before using these systems for research evaluation, and significant use should support rather than replace human expert judgement.

  • Takeaways & Limitations

    The extent of bias is unclear because the study lacks expert quality scores and cannot determine which compared approach is biased.

Abstract

from arXiv · show

Large Language Models (LLMs) can estimate the quality of published journal articles, potentially supporting human assessment when evaluations are needed. Whilst there are reasons to believe that LLMs may have biases in this role, there is no statistically strong evidence yet. The current article addresses this gap with an exploration of the types of articles that attract high or low LLM scores in 73,489 articles from 15 health and life sciences journals. Based on comparing the words in the titles and abstracts of higher and lower scoring articles for two LLMs in various ways, the results suggest that topics favoured by GPT-OSS-120B include viruses, genes and cells and its disfavoured topics include surveys, patients and students. It is not clear whether these patterns reflect underlying quality differences or AI biases, however. The same method found systematic differences between the topics favoured by GPT-OSS-120B and Gemma 3 27B, such as Gemma 3 27B giving relatively higher scores for machine learning research, proving that at least one of the two LLMs has AI bias. Finally, comparing the scores for full-text articles compared to scores for titles and abstracts also finds differences for both LLMs, showing that they both can exhibit AI bias for at least one of these two input types, and probably both. Overall, the results show that it is important to consider LLM biases when deciding whether to use them for research evaluation tasks.

1 Introduction

Research quality evaluation increasingly relies on assessments that can be costly or difficult for experts, creating interest in LLM-based scoring. The study addresses limited evidence about LLM bias by testing whether article types receive systematically different scores across health and life sciences.

  • Motivation: Senior researchers evaluate articles for hiring, promotion, funding, departmental monitoring, and national assessment, but may lack time or specialist expertise.This pressure can encourage shortcuts such as journal, author, citation, or impact-factor indicators.
  • Motivation: LLM judgements offer an alternative because, with careful prompting, their article scores correlate weakly to moderately with expert judgements.The passage presents this as a potential support for human assessment, not a replacement for it.
  • Research gap: Prior research suggested possible LLM bias in research-quality evaluation, but had not provided statistically strong proof.Earlier studies often used ChatGPT-derived scores from titles and abstracts rather than full texts.
  • Research gap: The study identifies three gaps: possible journal-style confounding, limited coverage beyond clinical medicine, and limited evidence from full-text inputs.These gaps motivate comparisons within and across journals, across health and life sciences, and between full-text and title/abstract inputs.
  • Research questions: The research questions test whether article types receive different LLM scores, whether patterns differ between LLMs, and whether they differ by input type.The comparisons are defined both within and across journals.

2 Methods

The study scored randomly selected articles from large health and life sciences journals with two LLMs using title/abstract and full-text inputs. It then used word-frequency tests and descriptive clustering and phrase analyses to examine systematic score differences.

  • Study design: The study obtained full-text articles from selected large health and life sciences journals and scored them with two LLMs using two input formats.The inputs were title/abstract and full text, with the latter including the title, abstract, and main body without references.
  • LLMs and prompts: Gemma 3 27B and GPT-OSS-120B were selected as open-weight models and prompted with REF2021 quality criteria.The criteria were rigour, originality, and significance, with scores from 1* to 4*; the newer GPT-OSS-120B was foregrounded.
  • Scoring: Each article was submitted five times for each prompt, and the average of the five overall scores became its final score.Repeated submissions addressed variation between identical prompt runs.
  • Statistical analysis: Word-frequency 2 x 2 chi-squared tests compared term occurrence in higher- and lower-scoring articles, using the median score as the cutoff.Benjamini-Hochberg correction was applied because separate tests were conducted for many words.
  • Descriptive analyses: K-means clustering grouped articles using TF-IDF features from titles and abstracts, while phrase analyses summarized mean scores for terms occurring in at least 25 articles.The clustering used 25 clusters and combined words, two-word phrases, and three-word phrases; both analyses were described as not statistically robust.

3 Results

The study covered 15 large health and life sciences journals, with journal-level scores strongly correlated across LLMs and input types. The reported score sets therefore showed substantial agreement at the journal-average level.

  • Journal coverage: 15 large health and life sciences journals were selected, spanning broad areas from health systems and whole-organism research to small-scale analyses.All except one published at least 5000 articles during 2020–2023.
  • Journal-level results: Journal-average scores correlated strongly across the four LLM-and-input combinations, ranging from r = 0.93 to r = 0.99.The lowest correlation was GPT-OSS-120B titles/abstracts versus Gemma 3 full texts; the highest was GPT-OSS-120B titles/abstracts versus GPT-OSS-120B full texts.
  • Journal-level results: Table 1 reports the selected journals, randomly selected article counts from 2020–2023, and mean scores from five runs for each stated LLM input.The table’s score summaries are organized around journal sampling and the specified processing condition.

3.1 Overall differences

Across the 15 journals, GPT-OSS-120B associated higher scores with molecular and theoretical topics and lower scores with surveys, patients, and students. Linguistic style also mattered, while descriptive analyses showed substantial score differences across phrases and topic clusters.

  • Viruses, cells, genes, genomics, mice, and proteins occurred disproportionately more often in higher-scoring articles.
  • Questionnaires, statistics, medical and hospital research, patients, and students occurred disproportionately more often in lower-scoring articles.
  • First-person plural writing, including phrases such as “Here we show…”, was the strongest overall linguistic pattern associated with score differences.The word “we” remained common in both higher- and lower-scoring articles, whose score distributions were similar.
  • 3.1.1 Descriptive analysis of overall differences: “Cryo-EM” had the highest average phrase score, whereas “dental students” and “Iasi” were among the lowest-scoring phrases.Phrase comparisons included only phrases appearing in at least 25 articles and used the overall mean as a reference.
  • 3.1.1 Descriptive analysis of overall differences: Virus studies formed the highest-scoring topic cluster, while Covid-19 and health studies scored lowest; molecular and theoretical studies exceeded human-level studies.

3.2 Within-journal differences

Within-journal analyses found that GPT-OSS-120B’s associations with higher and lower scores persisted after controlling for differences between journals. The results included both general stylistic terms and specialist topics.

  • Terms associated with higher GPT-OSS-120B scores appeared in multiple journals, showing that the overall patterns were not entirely attributable to journal differences.Within-journal comparisons address the possibility that journals differ in average research quality.
  • Higher-scoring terms included specialist topics such as Arabidopsis, the rockcress model plant organism.
  • Terms associated with lower GPT-OSS-120B scores were also identified separately for each journal.

3.3 Differences for the journal Cancers

The Cancers analysis used repeated scoring and clustering to examine systematic topic differences within one journal. Many cluster-score differences were statistically significant, with more theoretical and smaller-scale research scoring higher.

  • For Cancers, each article received 30 scores to improve score quality despite the smaller sample size.The additional scores were averaged into the final overall score, and this dataset was used only for Figure 5.
  • Many differences between Cancers topic-cluster average scores were statistically significant.
  • More theoretical and smaller-scale research scored higher in the Cancers topic-cluster comparison.

3.4 RQ2: Comparison with Gemma 3 27B

Gemma 3 27B and GPT-OSS-120B associate broadly similar topics with high and low article scores, but systematic differences show that their topic preferences are not identical.

  • Descriptive comparison: “Gut” ranked 140th among Gemma 3 27B’s high-score terms but did not rank in GPT-OSS-120B’s top 1,000.This illustrates that similar overall patterns can conceal substantial term-level differences.
  • Descriptive comparison: High- and low-scoring phrases for Gemma 3 27B substantially, but only partially, overlap with those for GPT-OSS-120B.Their overall topic hierarchies are also similar, suggesting broadly similar results.
  • Between-model differences: GPT-OSS-120B relatively favours virus and genome research, while Gemma 3 27B relatively favours machine learning and randomised trial research.GPT-OSS-120B also statistically favours titles and abstracts mentioning China, particularly Chinese province names.

3.5 RQ3: Full text against title/abstract

Both LLMs produce systematically different topic-associated scores depending on whether they receive full texts or only titles and abstracts, confirming input-dependent score patterns.

  • GPT-OSS-120B: Articles mentioning patients were more likely to receive lower GPT-OSS-120B scores from full text than from titles and abstracts.The comparison identifies systematic term differences between the two input types.
  • Both models: The study reports systematic title/abstract-versus-full-text term differences for both GPT-OSS-120B and Gemma 3 27B.Separate term tables identify articles scoring relatively higher with titles and abstracts or with full texts for each model.
  • GPT-OSS-120B: GPT-OSS-120B scores deep-learning and other machine-learning articles relatively higher from titles and abstracts, but food-quality articles relatively lower.These topic patterns are shown in the title/abstract-versus-full-text cluster comparison.
  • Gemma 3 27B: Gemma 3 27B also shows topic-dependent score differences between title/abstract and full-text inputs.Its cluster comparison is presented separately from GPT-OSS-120B’s.

4 Discussion

The discussion interprets the findings as evidence of systematic LLM score differences across topics, models, and input types, while stressing that their bias relative to expert judgment remains unresolved.

  • Limitations: The analysis is limited to selected journals, years, and two LLMs, and its results also depend on prompts and the UK REF quality definition.Other publishers, models, prompts, or quality criteria may produce different patterns.
  • Limitations: The study’s interpretation is limited because it lacks expert review scores as a quality gold standard.Consequently, pairwise differences show that at least one approach is biased, but not whether one or both are biased.
  • Limitations: The statistical tests partly violate independence because authors, teams, departments, and temporary trends can generate related articles.This dependence affects the assumptions underlying the statistical analyses.
  • RQ1: Word-frequency tests provide strong evidence that article types receive different scores within and across journals, but this does not by itself prove LLM bias.Topic-based disparities could reflect differences in expert evaluations or perceived importance rather than model-specific bias.
  • RQ2: GPT-OSS-120B relatively favours virus research, whereas Gemma 3 27B relatively favours machine-learning research.These statistically significant topic differences support a positive answer to RQ2, although no human expert gold standard identifies which model is less biased.
  • Interpretation: Different training data and human-feedback stages may contribute to the models’ divergent topic preferences.The paper proposes familiarity with differently represented topics and differing emphasis on dimensions such as rigour as possible explanations.
  • RQ3: Full-text and title/abstract inputs produce different topic-associated scores, supporting a positive answer to RQ3.The paper notes that full-text scores are generally higher, but the reason for the disparities is not obvious.

5 Conclusions

The study finds statistically significant, systematic topic and input-related differences in LLM research-quality scores, but cannot determine which model or input is least biased without expert-quality comparisons.

  • Conclusions: LLMs can show systematic topic and possibly linguistic-style biases in scores for health and life science journal articles.The study presents this as the first statistically significant evidence of systematic biases in LLM research-quality scores.
  • Conclusions: Full-text scoring changes the nature of score bias relative to title-and-abstract scoring, even though prior work found similar accuracy for the two inputs.Title-and-abstract inputs had tended to align more closely with human experts, but neither input is established as least biased here.
  • Implications: LLMs should support rather than replace human expert judgment for scoring published research, with bias identification and mitigation when used substantially.The paper notes that varying LLMs may not remove common biases between models.
Loading 2609.00323v1…