Source-linked AI summary
Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers
Yunhan Yang, Mike Thelwall, Guoxiu He
TL;DR
The paper addresses the limited understanding of how scientific critique functions differ between human and LLM peer reviews. Using ICLR 2025 data, it compares point-level weakness and question units under baseline and expert prompts, finding differentiated profiles rather than uniform human-likeness. Human reviews emphasize scientific framing and revision guidance, whereas LLM reviews emphasize explanation, integration, and formal argumentation.
Problem
Research has not systematically explained how scientific critique functions are distributed across human and LLM-generated reviews.
Method
The study compares ICLR 2025 human reviews with baseline- and expert-prompt LLM reviews using five theory-guided frameworks applied to weakness and question points.
Results
Human reviews foreground scientific framing and revision priorities, while LLM reviews foreground explanation, integration, and explicit argumentation.
Takeaways & Limitations
LLM-assisted peer review should expose its functional profile while retaining human experts for significance, revision prioritization, and accountable evaluation.
Takeaways & Limitations
The evidence is limited to ICLR 2025 machine-learning reviews and mainly Gemma3-27B-generated reviews, so other disciplines and model families require testing.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique they perform. Using ICLR 2025 peer-review data, we compare human reviews with LLM reviews generated under baseline and expert prompts. We operationalize scientific critique through two review acts, weakness critique and scientific questioning, and annotate point-level review text using five theory-guided frameworks: Anderson's knowledge types, Toulmin's argumentation model, Graesser's question depth, SOLO cognitive complexity, and Hattie's feedback functions. The results reveal a differentiated critique profile. Human reviews placed greater emphasis on scientific framing and revision guidance, more often identifying higher-order weaknesses and asking questions oriented toward improvement. LLM reviews showed higher rates of explanatory depth, integrative reasoning, and explicit argument structuring. Expert prompting did not make LLM critique uniformly more human-like; it partially narrowed some gaps but mainly amplified LLM-specific tendencies toward integration and formal argumentation. These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.
1 Introduction
The study reframes LLM peer review around the functions of scientific critique rather than human-likeness, usefulness, or textual overlap. It compares human and LLM reviews across weakness diagnosis, questioning, and prompt conditions.
- Motivation: Peer-review research has examined LLM usefulness and overlap with human comments more than the functions distributed across their critiques.This leaves open whether LLMs identify scientifically consequential weaknesses and support revision in the same ways as humans.
- Analytical focus: The study separates review text into weakness critique and scientific questioning because they expose diagnostic and inquiry-oriented aspects of evaluation.Weaknesses identify problems and justify judgments, whereas questions ask authors to clarify, justify, extend, or revise.
- Research questions: The research asks how baseline LLM weaknesses and questions differ from humans and whether expert prompting narrows or amplifies those differences.The questions target diagnostic content, argumentative support, explanatory depth, integrative complexity, and revision-oriented feedback.
- Study design: It compares human reviews with baseline and expert-prompt LLM reviews using theory-guided, point-level analysis of critique functions.The study examines ICLR 2025 reviews and converts annotated points into interpretable comparisons.
2 Literature Review
Prior work shows that LLMs can produce useful review-like feedback, but it does not systematically explain how critique functions are distributed across human and LLM reviews. The study therefore models peer review as structured evaluative information across targets, support, questioning, and revision.
- LLMs in scholarly evaluation: Existing studies report useful or overlapping LLM feedback, but also limitations in novelty assessment, flaw identification, and balanced multidimensional evaluation.The literature establishes participation in review workflows without resolving how scientific critique is functionally organized.
- Dimensions of critique: Scientific critique evaluates research importance, originality, methods, interpretation, constructiveness, and substantiation rather than a single quality dimension.Review comments address diverse targets including methodology, theory, writing, originality, impact, clarity, soundness, and recommendations.
- Questioning: Reviewer questions can test relationships among claims, evidence, assumptions, and possible revisions, initiating clarification, defense, or revision.Questioning is treated as a substantive form of scientific inquiry rather than a request for missing information alone.
- Analytical framework: Five theory-guided frameworks translate critique into indicators of content target, argumentative support, explanatory depth, integrative complexity, and feedback function.Anderson and Toulmin analyze weaknesses; Graesser, SOLO, and Hattie analyze scientific questions.
- Argumentative support: Critique also depends on how weaknesses are supported through evidence, reasoning, and relations among claims, data, and warrants.This distinguishes merely asserted criticism from arguments authors can evaluate and use.
3 Methods
The method constructs comparable human and LLM review-section text, decomposes it into weakness and question points, and applies five theory-guided frameworks. The resulting labels are aggregated into paper-level critique metrics.
- Text construction: Human review sections and LLM-generated reviews from full papers are decomposed into weakness and question points for comparable analysis.LLM reviews are produced under baseline and expert prompts.
- Theory-guided annotation: Anderson and Toulmin analyze weakness critique, while Graesser, SOLO, and Hattie analyze scientific questions.Together, the frameworks represent content targets, argumentative support, explanatory depth, integrative complexity, and feedback functions.
- Metric construction: Point-level annotations are aggregated into paper-level metrics comparing human, baseline-prompt LLM, and expert-prompt LLM critique profiles.Figure 1 presents the overall methodological framework linking these stages to the study output.
3.1 Data Construction
The data come from structured ICLR 2025 review reports and matched LLM-generated reviews. The main comparison uses the full corpus, while expert-prompt analysis uses a balanced 600-paper sample across submission outcomes.
- Empirical setting: ICLR 2025 was selected for its large machine-learning venue, structured open reviews, standardized sections, and accessible paper-level metadata.Submission types included accepted, rejected, withdrawn, and a small number of desk-rejected papers.
- Review construction: Human weakness and question sections were extracted from review pairs, while LLMs generated corresponding text from full papers under baseline and expert prompts.The expert prompt emphasized importance, contribution, claim-method-evidence alignment, alternative explanations, and robustness.
- Source data: LLM inputs were parsed from OpenReview submission PDFs available at the time of data collection.They were not taken from external arXiv records or separate proceedings files.
- Sampling: The expert-prompt experiment sampled 600 papers evenly across accepted, rejected, and withdrawn submissions.The sample contained 200 papers from each outcome group and served as a prompt-sensitivity analysis rather than a second full-corpus estimate.
3.2 LLM Configuration and Prompting
The study separates LLM review generation from annotation and uses local, reproducible model execution, with Gemma3-27B as the main model and Qwen3-32B for robustness checks.
- LLMs generated review text from full papers and annotated point-level review text, separating the analyzed critique from its measurement procedure.
- All models ran locally on a university high-performance computing cluster rather than through commercial APIs.This provided a consistent execution environment and avoided sending paper or review text to external services.
- Gemma3-27B served as the main generation and annotation model, while Qwen3-32B tested model-family sensitivity.The robustness design crossed generation and annotation models in four combinations.
3.3 Review Point Annotation and Validation
Review points were classified with five theory-guided schemes that capture critique content, argument structure, question complexity, and feedback function, with Toulmin features retained as separate binary annotations.
- Weakness points were coded with Anderson’s knowledge types and Toulmin’s argumentation model, while questions used Graesser, SOLO, and Hattie frameworks.These schemes measure critique content, argumentative structure, explanatory depth, integrative complexity, and feedback function.
- Toulmin annotations included claim, data, warrant, backing, qualifier, and rebuttal, with summary categories from claim-only to dialectical critique.The elements were coded separately because multiple elements could co-occur within one weakness point.
- The five-scheme classification approach was summarized in Table 2, while Table 3 reported agreement between human audit labels and LLM annotations using Qwen3 kappa.
- Anderson categories distinguished factual, procedural, conceptual, and metacognitive critique targets.
- Question annotations represented cognitive depth from shallow to deep, SOLO complexity from pre-structural to extended abstract, and feedback orientation through task, process, and feed-forward functions.
3.4 Paper-level Metrics
The paper aggregates point-level annotations into matched paper-level metrics for weakness critique and scientific questioning, covering critique level, argumentation, question depth, complexity, and revision guidance.
- Annotations were aggregated separately for weakness points and question points by paper and source, enabling matched comparisons.
- Weakness critique metrics: High-order critique rate measured the share of weakness points coded as conceptual or metacognitive.
- Weakness critique metrics: Argument depth score averaged each weakness point’s highest support-related Toulmin element, from claim-only through rebuttal.Qualifiers were excluded because they indicate uncertainty calibration rather than argumentative elaboration.
- Weakness critique metrics: Expert-like weakness rate captured weaknesses that were both high-level in content and sufficiently argued.This required conceptual or metacognitive content plus reasoned, contextualized, or dialectical argumentation.
- Scientific questioning metrics: Question metrics included deep-question rate, SOLO complexity score, and feed-forward rate, alongside auxiliary depth, relational, task, process, mechanistic, and revision-oriented measures.
3.5 Statistical Comparisons and Prompt Effects
The analysis compares human and baseline-LLM critique at matched-paper level and tests whether expert prompting reduces human–LLM metric gaps or shifts critique functions in another direction.
- Human and baseline-prompt LLM values were compared at matched-paper level using paired differences, rank-biserial effect sizes, and Wilcoxon signed-rank tests.Comparisons were also repeated across accepted, rejected, and withdrawn submissions.
- The study assessed whether expert prompting moved LLM-generated review profiles closer to human profiles on the same paper-level metrics.This comparison used a stratified 600-paper sample containing accepted, rejected, and withdrawn submissions.
- Prompt-induced convergence was defined so positive values indicated movement toward the human profile and negative values indicated movement farther away.
- Directional prompt shifts identified which critique functions expert prompting amplified or reduced.
3.6 Prompt Convergence Metrics for Prompt Effects
The study uses four composite metrics to distinguish expert-prompt effects that move LLM critique toward the human baseline from shifts that amplify formal, expert-like features without increasing human-like convergence.
- Composite metric design: Four composites combine labels from Anderson, Toulmin, Graesser, SOLO, and Hattie to assess prompt effects on weakness critique and questioning.The metrics separate substantive convergence from formal amplification.
- Weakness critique: Supported high-order critique measures weakness points that are conceptual or metacognitive and include data, warrant, backing, or rebuttal.It captures the overall prevalence of high-level critique grounded in explicit support.
- Weakness critique: Support share within high-order critique measures whether high-order weaknesses contain at least one support element.This distinguishes grounded high-level critique from critique that remains merely abstract.
- Scientific questioning: Generic deep question rate measures questions that are deep under Graesser but not feed-forward under Hattie.It captures explanatory or mechanistic questioning without direct revision guidance.
- Scientific questioning: Substantive revision-oriented deep question rate measures questions that are deep, feed-forward, and relational or extended abstract under SOLO.This composite combines explanatory depth, integrative complexity, and revision guidance.
- Interpretation: All four metrics are interpreted against the human baseline rather than as monotonic quality scores.An increase with a smaller human–LLM gap indicates substantive convergence; an increase with a larger gap indicates formal amplification.
4 Results
Human and baseline-prompt Gemma reviews differed systematically in what they criticized, how they argued weaknesses, and how they used questions for explanation versus revision guidance. Expert prompting selectively shifted these profiles but generally amplified Gemma’s integrative and formal tendencies rather than producing broad human-like convergence.
- Weakness Critique: Human reviews had a higher high-order critique rate, while Gemma reviews had greater Toulmin argument depth.High-order critique was 0.451 for humans versus 0.379 for Gemma; argument depth was 1.581 versus 1.752, respectively.
- Weakness Critique: The integrated expert-like weakness metric was slightly higher for Gemma, apparently because argumentative development outweighed its lower high-order critique rate.The metric was 0.091 for humans and 0.115 for Gemma; the passage attributes the difference mainly to argument depth.
- Weakness Critique: Gemma weaknesses were more procedural and warrant-based, whereas human weaknesses contained more conceptual and metacognitive critique.The Gemma advantage in argument depth was primarily driven by reasoned critique, especially warrant-based statements.
- Weakness Critique: These weakness-profile differences remained stable across accepted, rejected, and withdrawn submissions.Humans consistently showed higher high-order critique, while Gemma consistently showed higher argument depth and expert-like weakness rates.
- Scientific Questioning: Gemma reviews contained deeper and more relational questions, while human reviews more often used questions as actionable feed-forward guidance.Gemma’s deep-question rate differed by +0.056 and relational-question rate by +0.135; human feed-forward rate differed by -0.059 when calculated as Gemma minus human.
- Scientific Questioning: Questioning differences also persisted across submission outcomes, indicating a stable difference in questioning profile rather than an outcome-specific pattern.Stratified effects closely tracked overall estimates across accepted, rejected, and withdrawn papers.
- Expert Prompting: Expert prompting selectively changed Gemma’s profile but mainly increased formal weakness development and integrative or process-oriented questioning.It moved argument depth from 1.766 at baseline to 2.961 under the expert prompt, while feed-forward rate fell from 0.364 to 0.307.
- Expert Prompting: Expert prompting did not produce broad human-like convergence in either critique selection or revision-oriented questioning.It increased supported high-order critique and generic deep questioning but did not clearly improve revision-oriented scientific inquiry.
5 Discussion
The study reframes differences between human and LLM peer reviews as differences in how critique functions are distributed, rather than as a single overall-quality gap. Human reviews emphasized scientific priorities and revision guidance, while LLM reviews emphasized explanation, integration, and formal argumentation; expert prompting intensified rather than erased this distinction.
- Discussion: The study maps how scientific critique is functionally organized instead of evaluating LLM reviews primarily by usefulness or human-likeness.It treats differences as a redistribution of evaluative work across review text.
- Human critique: Human reviews more often identified higher-order weaknesses and directed questions toward concrete improvement.They more often expressed what was scientifically at stake and what authors should revise.
- LLM critique: Baseline-prompt LLM reviews more often contained explanatory questions, integrative complexity, and explicitly structured arguments.These tendencies can broaden possible questions, make warrants explicit, and connect multiple study elements.
- Prompting: Expert prompting partially narrowed some gaps but mainly intensified LLM tendencies toward integrative reasoning and formal argument structuring.It also produced less revision guidance, so more systematic or authoritative language did not necessarily indicate improved evaluative judgment.
- Implications: The findings support cautious complementarity: LLMs may broaden issues considered, while human experts judge significance, prioritize revisions, and make accountable decisions.Systems should expose where LLMs amplify or over-formalize critique rather than make them indistinguishable from human reviewers.
- Limitations: The evidence is bounded by ICLR 2025, primarily Gemma3-27B reviews, and function-rate analysis that does not show an LLM cannot produce less-represented functions.Generalizability requires other disciplines, review cultures, model families, and downstream outcome measures.
6 Conclusion
The study concludes that human and LLM reviews expose different aspects of scientific critique. LLMs should therefore be assessed by the critique functions they foreground and the human judgment needed to interpret and govern those differences.
- 6 Conclusion: Separating weakness critique from scientific questioning shows that human and LLM reviews make different aspects of scientific critique visible.The study treats LLM-generated peer review as a redistribution of evaluative work rather than a direct approximation of human review.
- 6 Conclusion: Human reviews foregrounded scientific framing and revision priorities, whereas LLM reviews foregrounded explanation, integration, and explicit argumentation.Expert prompting changed this distribution without making LLM critique uniformly more human-like.
- 6 Conclusion: LLMs should not be evaluated only by overall similarity to human reviewers.Their role depends on which critique functions become more or less visible and on the human judgment required to interpret and govern those differences.