Source-linked AI summary

Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal

Timo van der Kuil, Bruno Messina Coimbra, Mirjam van Zuiden, Robert A. Bagheri, Rens van de Schoot, Klaas Dieleman, Berend Greijn, Stefan Houkes, Sebastiaan Rodenhuis, Elizabeth M. Grandfield

arXiv:2608.20385v1cs.CL

TL;DR

Checklist-based quality appraisal is time-consuming and sensitive to ambiguous criteria, while the effect of checklist design on LLM agreement with experts remains unclear. The paper compares LLM and expert assessments with GRoLTS across three domains and two checklist versions, finding that disagreement patterns can guide checklist revision and that revised items improve agreement.

  • Problem

    Quality appraisal in systematic reviews is labor intensive, and it remains unclear how checklist design affects LLM agreement with expert judgments.

  • Method

    The study applies a RAG-based GRoLTS annotation pipeline and compares LLM and expert judgments across three research topics, checklist versions, and agreement measures.

  • Results

    Performance varies across checklist items, ambiguous and conditional criteria show the greatest disagreement, and revising them improves raw and chance-corrected agreement.

  • Takeaways & Limitations

    Analyzing human–LLM disagreement can identify problematic checklist items and support iterative improvement of LLM-assisted research-synthesis workflows.

Abstract

from arXiv · show

Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks, appraisal checklists are typically treated as fixed inputs, and it remains unclear how their design affects agreement with expert judgments. Therefore, we investigate (1) whether LLMs can approximate human judgments in checklist-based appraisal and (2) whether patterns of human-LLM disagreement can be used to identify and improve ambiguous checklist items. Using the Guidelines for Reporting on Latent Trajectory Studies (GRoLTS) checklist, we compare LLM-generated assessments with expert annotations across three research topics and two checklist versions. Agreement is assessed using item-level accuracy, chance-corrected agreement, and preservation of study-level rank ordering. We find that performance varies substantially across checklist items, with ambiguous and conditional criteria producing the greatest disagreement. Revising these items improves both raw and chance-corrected agreement. Although item-level misclassifications persist, LLM-generated scores often preserve the relative ranking of studies when high-agreement items are retained. These results indicate that reliable LLM-assisted appraisal depends not only on model choice but also on checklist design. The findings suggest that analyzing human-LLM disagreement can help identify problematic checklist items and support the iterative improvement of research synthesis workflows.

1 Introduction

The paper examines whether LLMs can support checklist-based quality appraisal and whether human–LLM disagreements reveal checklist items needing revision. It evaluates agreement, checklist design, and study-ranking preservation across multiple domains and checklist versions.

  • Motivation: Quality appraisal is labor intensive, cognitively demanding, and can take 30 to 60 minutes per study.The process may also be susceptible to fatigue-related errors in large reviews.
  • Motivation: Checklist-based appraisal is difficult to automate because criteria often require interpreting implicit assumptions and domain-specific judgments.Human reviewers also frequently disagree when checklist wording is ambiguous or underspecified.
  • Study design: The study compares item-level human–LLM judgments using a RAG pipeline applied to the GRoLTS reporting checklist.The checklist targets transparency and consistency in studies using latent trajectory models.
  • Study design: The authors revise the checklist to reduce ambiguity, simplify complex or conditional items, and improve interpretability for human and automated annotators.The revised checklist is tested on PTSD, adolescent delinquency, and educational achievement datasets.
  • Evaluation: Agreement is assessed through item-level accuracy, chance-corrected metrics, and preservation of study-level reporting-quality rankings.These analyses identify reliably automatable items, evaluate checklist-design effects, and assess rank preservation for evidence synthesis.

2 Methods

The methods combine GRoLTS checklist appraisal across three topic datasets with a revised checklist and a retrieval-augmented annotation pipeline. Full-text articles are segmented, embedded, retrieved by checklist-item similarity, and supplied to LLMs for structured binary judgments.

  • Datasets and checklist: GRoLTS contains 21 binary yes-or-no items, with total yes responses summarizing reporting quality.The checklist was originally developed for human experts rather than automated use.
  • Datasets and checklist: Three datasets cover PTSD, educational achievement, and adolescent delinquency studies using LGMM or LCGA.The PTSD dataset contains 38 previously annotated studies, while the educational-achievement and adolescent-delinquency datasets were newly retrieved and screened.
  • Checklist revision: Checklist version 2 was developed from preliminary LLM outputs, human–LLM disagreement analysis, and expert feedback.Revisions reduce ambiguity, remove double-barreled and conditional sub-items, and align criteria with explicitly reported information.
  • LLM annotation pipeline: The five-stage pipeline acquires and preprocesses articles, segments text, embeds chunks and checklist items, retrieves relevant chunks, and generates binary judgments.Retrieved evidence is incorporated into a structured prompt containing a rationale, supporting quotation, and YES/NO answer.
  • LLM annotation pipeline: Articles are converted from PDF to Markdown, preserving structural elements such as headings and tables for retrieval and interpretation.Documents are divided into 1,000-word chunks with 50-word overlap.
  • LLM annotation pipeline: Qwen3-Embedding-8B maps text chunks and checklist items into a shared m-dimensional vector space for semantic matching.For each document–question pair, the ten highest-similarity chunks are selected as contextual input.

3 Results

Human–LLM agreement varied substantially across GRoLTS items, and revisions targeting ambiguous criteria improved agreement across checklist versions and domains. Rank ordering was best preserved when study scores retained high-agreement items.

  • Human–LLM Agreement: GRoLTS v1: Mean item accuracy ranged from approximately 0.10 for question 18 to nearly 1.0, showing strong question-level heterogeneity in GRoLTS v1.Lower-accuracy questions also showed greater variability across LLMs, while higher-accuracy questions clustered more tightly.
  • Item-level Challenges and Checklist Revisions: Low-accuracy or high-variability items revealed disagreement from partial-versus-complete reporting requirements and broad phrasing.These recurring patterns guided targeted checklist revisions.
  • Item-level Challenges and Checklist Revisions: Revisions clarified criteria, distinguished partial from complete reporting, and removed or reformulated persistently ambiguous items to improve alignment.Examples included splitting reporting requirements, omitting ambiguous item 18, and adding explicit distributional examples.
  • Human–LLM Agreement: GRoLTS v2: v2 mean accuracies generally rose from v1’s approximately 0.10–0.20 low range to roughly 0.40–0.60, with more items in the 0.7–1.0 range.Variability also decreased for many mid- and high-performing items, although some question-level challenges remained.
  • Human–LLM Agreement: GRoLTS v2: Agreement patterns were similar across Achievement, Delinquency, and PTSD for higher-ranked items, while lower-ranked items showed greater divergence.PTSD v2 deviations may reflect remapping artifacts or differences in human interpretation of revised items.
  • Chance-Corrected Agreement: Cohen’s κ ranged from 0.41 to 0.72 for v2 versus 0.31–0.52 for PTSD v1, while Fleiss’ κ was approximately 0.60–0.61 in v2 versus 0.41 in v1.GPT-5 mini achieved the highest human–LLM agreement among the v2 models, with relatively small differences among the remaining models.
  • Rank-order Consistency: Overall rank correlations ranged from 0.36 to 0.72, with high-agreement items producing ρ ≈0.58–0.93 and low-agreement items sometimes approaching zero or becoming negative.These results support using LLMs for well-specified criteria while reserving ambiguous or conceptually complex items for human review.

4 Discussion

LLM-assisted appraisal showed moderate to substantial agreement with human experts, but performance depended strongly on checklist-item clarity and structure. Disagreement patterns guided revisions that improved agreement and preserved study rankings on high-agreement items, while limitations remained around prompting, human variability, dataset scope, and pipeline optimization.

  • Moderate to substantial agreement was observed between LLM-generated and human annotations across case studies and checklist versions.
  • 4.1 Question-Level Agreement as a Diagnostic Tool: Question-level performance varied substantially, with conditional phrasing, ambiguous scope, and broadly defined criteria producing the most disagreement.
  • 4.1 Question-Level Agreement as a Diagnostic Tool: Analyzing disagreement helped identify ambiguous or underspecified items and guided revisions that improved raw and chance-corrected agreement.
  • 4.2 Implications for LLM-Assisted Quality Appraisal: LLMs often preserved the relative ordering of studies despite item-level misclassifications, especially when high-agreement items were retained.
  • 4.2 Implications for LLM-Assisted Quality Appraisal: Model differences were present but comparatively modest, indicating that retrieved information and checklist design also influenced performance.
  • 4.3 Limitations and Future Work: The study was limited by prompt sensitivity, human inter-rater variability, evaluation across few datasets, and pragmatic rather than exhaustive RAG optimization.
  • 4.3 Limitations and Future Work: Future work should examine iterative human–LLM refinement across additional domains and larger corpora.

Declarations

The paper reports funding, ethics, consent, data-availability, code-availability, and author-contribution information.

  • The work was supported by the Dutch Research Council and the Dutch national e-infrastructure with SURF Cooperative support.
  • Ethics approval, consent to participate, and consent for publication were each reported as not applicable.
  • Data and code were made available through OSF and GitHub, respectively.
  • Author contributions covered conceptualization, methodology, analysis, investigation, software, data curation, resources, visualization, administration, supervision, funding, and writing.

A Original GRoLTS Questions (v1)

The original GRoLTS v1 questions cover reporting of time metrics, missing data, software, model specifications, model comparisons, class characteristics, plots, and syntax availability.

  • The checklist asks whether the metric or unit of time used in the statistical model is reported.
  • It asks whether mean and variance of time within a wave are presented.
  • Missing-data items address the missing-data mechanism, variables related to attrition, and how missing data were handled.
  • The checklist asks whether information about observed-variable distributions and the software is reported.
  • Additional items cover alternative within-class and between-class specifications, covariate replicability, random starts, final iterations, and statistical model-comparison tools.
  • It asks whether total fitted models, cases per class, entropy, estimated mean trajectory plots, final-solution characteristics, and syntax files are reported.

B Revised GRoLTS Questions (v2)

The revised GRoLTS questions cover reporting of statistical modeling, data handling, model specification, class solutions, trajectories, and reproducibility details.

  • The checklist asks whether the metric or unit of time used in the statistical model is reported.
  • Items address missing-data handling, observed-variable distributions, software, within-class heterogeneity, trajectory forms, covariate reproducibility, and optimization settings.
  • The checklist evaluates reporting of model-comparison tools, fitted-model counts, one-class solutions, class sizes, and entropy.
  • Further items assess plots of estimated mean trajectories, observed individual trajectories by latent class, numerical class-solution characteristics, and syntax-file availability.

C Prompt Template

The prompt template instructs an LLM to answer each checklist question from retrieved paper context in a fixed, evidence-oriented format.

  • The task asks whether an academic paper reports specified methodological or statistical information.
  • The model must use only the supplied markdown-formatted context from a single academic paper.
  • Each response contains a brief rationale, verbatim supporting evidence, and a final binary YES/NO judgment.
  • The rules require explicit evidence, prohibit outside knowledge, and require NO when information is missing, unclear, or only implied.
  • The prompt combines the checklist question with retrieved context as the model input.

D Item-level accuracy per LLM for GRoLTS v1 on PTSD

Figure 5 presents item-level agreement between five LLMs and human labels for the GRoLTS v1 checklist in the PTSD use case.

  • Figure 5 organizes Q1–Q21 by rows and five LLMs by columns, with each cell showing agreement with human labels.
  • The figure reports mean accuracy per model, mean accuracy per question, and the proportion of positive labels for each item.

E Item-level accuracy per LLM for GRoLTS v2 on Achievement, Delinquency, and PTSD

Figures 6–8 report item-level agreement for the GRoLTS v2 checklist across Educational Achievement, Adolescent Delinquency, and PTSD use cases.

  • Educational Achievement: Figure 6 evaluates five LLMs on GRoLTS v2 items Q1–Q19 for the Educational Achievement use case.
  • Adolescent Delinquency: Figure 7 evaluates five LLMs on GRoLTS v2 items Q1–Q19 for the Adolescent Delinquency use case.
  • PTSD: Figure 8 evaluates five LLMs on GRoLTS v2 items Q1–Q19 for the PTSD use case.
Loading 2608.20385v1…