Source-linked AI summary

Perspectives on Large Language Models for Relevance Judgment

Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, Henning Wachsmuth

arXiv:2304.09161v2cs.IRcs.CY

TL;DR

IR evaluation traditionally relies on costly human relevance judgments, creating interest in whether LLMs can provide reliable assistance or automation. This perspectives paper organizes options along a human–machine collaboration spectrum, reviews risks and opportunities, and pilots agreement with trained assessors. The authors find promise in mimicking human assessments, but conclude that current concerns prevent fully automated deployment while supporting further human–LLM collaboration.

  • Problem

    Human assessors are required for Cranfield-style relevance judgments, but the process is time-intensive and costly, while reliable automated alternatives remain unclear.

  • Method

    The paper synthesizes prior approaches, proposes a human–machine collaboration spectrum, and conducts a pilot comparison of LLM-generated judgments with trained human assessments.

  • Results

    LLMs show promise in mimicking human relevance assessments, but agreement varies by collection and relevance label.

  • Takeaways & Limitations

    LLMs may support human assessors across a collaboration spectrum, but the paper does not support fully automated relevance annotation at present.

  • Takeaways & Limitations

    There is currently no proof that LLM-generated evaluations have any relationship to reality beyond their ability to mimic human language.

Abstract

from arXiv · show

When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems. In this perspectives paper, we discuss possible ways for LLMs to support relevance judgments along with concerns and issues that arise. We devise a human--machine collaboration spectrum that allows to categorize different relevance judgment strategies, based on how much humans rely on machines. For the extreme point of "fully automated judgments", we further include a pilot experiment on whether LLM-based relevance judgments correlate with judgments from trained human assessors. We conclude the paper by providing opposing perspectives for and against the use of~LLMs for automatic relevance judgments, and a compromise perspective, informed by our analyses of the literature, our preliminary experimental evidence, and our experience as IR researchers.

1 INTRODUCTION

The paper examines whether LLMs can reduce the costly human effort required for relevance judgments while preserving reliable retrieval evaluation. It organizes possible uses along a human–machine collaboration spectrum and presents balanced perspectives informed by literature analysis and a pilot experiment.

  • 1 INTRODUCTION: Human relevance judgments are central to Cranfield-style IR evaluation but require substantial time and money.TREC-8 Ad Hoc alone involved more than 86,000 pooled documents, over 700 assessor hours, and about USD 15,000.
  • 1 INTRODUCTION: LLMs may assist or replace parts of relevance assessment, but their agreement with human judgments remains unclear.The paper focuses on text-based test collections to compare LLM and human assessors within the Cranfield paradigm.
  • 1 INTRODUCTION: The paper presents opposing perspectives for and against LLM assessors, plus a compromise perspective based on literature analysis, pilot evidence, and IR research experience.The supplied contribution passage is truncated after describing the perspectives as informed by these sources.
  • 1 INTRODUCTION: Prior approaches to scaling judgment collection include crowdsourcing, incomplete-judgment metrics, similarity-based label transfer, active learning, and automatic question-based evaluation.These methods aim to reduce assessment cost or target which documents and information units receive judgments.
  • 1 INTRODUCTION: The collaboration spectrum ranges from fully manual judgments to fully automatic judgments, with several intermediate forms of machine support.Intermediate options include highlighting, summaries, competence partitioning, multiple LLM suggestions, and human acceptance or rejection of machine judgments.

3 SPECTRUM OF HUMAN–MACHINE COLLABORATION

The paper proposes a spectrum describing how humans and LLMs divide relevance-judgment labor, from unaided human assessment to complete machine replacement. Intermediate levels retain different amounts of human control and verification.

  • 3 SPECTRUM OF HUMAN–MACHINE COLLABORATION: The spectrum spans manual human judgment, increasing LLM assistance, human verification, and fully automated assessment.Its endpoints are humans judging without LLM interaction and LLMs replacing humans completely.
  • AI Assistance: AI assistance can provide document summaries or automatically count information nuggets while humans make the relevance decision.These approaches compress documents or operationalize manually defined relevance information for more efficient assessment.
  • Human Verification: Human verification uses an LLM’s judgment and rationale as a suggestion, with people intervening especially on challenging or low-confidence cases.A preference-testing variant lets humans choose between two LLM judgments, while humans retain the ultimate decision when needed.
  • Fully Automated: Fully automated assessment would replace human judgment, but detecting whether LLMs surpass humans remains an open problem.Human-annotated gold standards cannot reveal superiority once human judgment becomes the measurement limit.
  • 3 SPECTRUM OF HUMAN–MACHINE COLLABORATION: The proposed ideal is competence partitioning, assigning tasks to humans or machines according to who is better suited at the best cost.The paper frames this as the central question for selecting a point on the collaboration spectrum.

4 OPEN ISSUES AND OPPORTUNITIES

The paper identifies unresolved questions about the cost, quality, bias, factuality, and explainability of LLM-supported relevance judgments. It also considers whether collaboration could enable broader and more realistic evaluations, while fully automated use remains constrained by uncertainty.

  • 4.1 LLM Judgment Cost and Quality: The effectiveness of LLM-based judgment support is uncertain because eventual LLM capabilities and the quality-cost trade-off are not yet established.The paper compares assessor types and judgment tasks but treats LLM replaceability as uncertain because development remains at an early stage.
  • 4.1 LLM Judgment Cost and Quality: Human feedback, fine-tuning, or active learning could align LLM suggestions more closely with individual assessors’ relevance judgments.An LLM could begin with mild suggestions and continuously learn from the assessor’s actual decisions.
  • 4.1 LLM Judgment Cost and Quality: Using multiple LLM assessors may not provide independent judgments because many models share similar training data and may share biases.Training or fine-tuning models on different data or user types is proposed as a way to obtain less correlated judgments.
  • 4.2 Human Verification: LLM relevance judgments must address factuality, misinformation, bias, hallucination, and the difficulty of translating nuanced relevance guidelines into prompts.Topical matching alone may be insufficient when documents are factually false, culturally specific, or stylistically unsuitable.
  • 4 OPEN ISSUES AND OPPORTUNITIES: Collaborative human–machine judgment might scale evaluations using more comprehensive and realistic notions of relevance than standard manual setups.Such setups can include changing information needs and other notions beyond simple topical relevance.

5 PRELIMINARY ASSESSMENT

The preliminary assessment compares GPT-3.5 and YouChat with human assessors across binary and graded relevance judgments on TREC-8 and TREC-DL 2021. Agreement varies by collection, model, and relevance grade, with weaker alignment on some difficult distinctions.

  • Methodology: The experiment compares GPT-3.5 and YouChat with human assessors on binary and graded judgments across TREC-8 and TREC-DL 2021.The study uses two test collections, two LLMs, two judgment types, and tailored prompts.
  • Methodology: The sampled judgments comprise 1000 topic–document pairs per collection, though YouChat experiments use 100 random samples per relevance grade.The restriction reflects YouChat’s limited scalability.
  • Methodology: The prompts were intentionally kept simple rather than optimized, establishing a first baseline for comparison.Prompt engineering was left for future work.
  • Results: On TREC-8, GPT-3.5 agrees with human judgments for 90% of human-labeled non-relevant documents but only 47% of relevant documents.YouChat shows the same divide, with 74% agreement for non-relevant and 33% for relevant documents.
  • Results: On TREC-DL 2021, YouChat agrees on 96 of 100 highly relevant question–passage pairs, while agreement for non-relevant pairs is 42 of 100.GPT-3.5 has problems with the middle grades 1 and 2.
  • Results: The authors hypothesize that subtle distinctions used by human assessors may not be recognizable to LLMs in binary and graded judgments.This interpretation concerns difficult relevance boundaries and middle grades.

6 RE-JUDGING TREC 2021 DEEP LEARNING

The paper re-judges the TREC-DL 2021 passage-ranking pool with GPT-3.5 and compares resulting agreement and system evaluations with official human judgments. The LLM judgments show fair binarized agreement, preserve the top-ranked run, and distinguish fewer run pairs.

  • 6.1 Methodology: The re-judging experiment applies GPT-3.5 judgments to the TREC-DL 2021 passage-ranking pool while following the track’s graded-judgment methodology.The original pool contains 10,828 judgments on a four-point relevance scale.
  • 6.1 Methodology: The prompt uses few-shot examples for all relevance levels, including two examples of non-relevant passages.The judgment cost was about USD 0.01 each, compared with USD 0.25 per human judgment in a similar task.
  • 6.2 Results: Cohen’s κ is 0.26 for binarized judgments, indicating fair agreement between GPT-3.5 and the official judgments.The binarization maps perfectly and highly relevant to relevant, and relevant and non-relevant to non-relevant.
  • 6.2 Results: The top run under official judgments remains the top run under LLM judgments when standard evaluation measures are computed.Figure 3 compares MAP and NDCG@10 effectiveness across submitted runs under both judgment sources.
  • 6.2 Results: LLM-based measures are less sensitive than human-based measures: 65% versus 72% of pairs under MAP, and 69% versus 74% under NDCG@10.Sensitivity is the proportion of submitted-run pairs distinguished by paired t-tests at p < 0.05.

7 PERSPECTIVES FOR THE FUTURE

The paper presents balanced perspectives on LLM-based relevance judgments, combining potential benefits with concerns about reliability, bias, and human verification. Its pilot evidence suggests promise, while favoring collaboration over unexamined full automation.

  • The paper frames LLM relevance judgment through opposing perspectives for and against automatic assessment, plus a compromise perspective.
  • LLMs can offer explanations, scalability, consistency, and some quality, making them potentially useful complements to human assessors.
  • Relevance is subjective and human-grounded, so LLM judgments require verification because their connection to reality remains unproven.
  • LLM judgments raise additional concerns about source attribution, prompt opacity, misinformation, intellectual property, and social bias.
  • The pilot found reasonable correlation with highly trained human assessors and similar leaderboards, supporting further study of automated judgments.
  • The paper identifies AI assistance as a credible but underexplored path and calls for research on humans verifying LLM suggestions and rationales.
  • The conclusion presents LLM relevance judgment as promising but not yet suitable for fully automated annotation, while proposing a spectrum of human support.

Laura Dietz University of New Hampshire

The introduction motivates LLM-assisted relevance judgment by contrasting costly human assessment with machines’ expanding role in IR. It proposes a collaboration spectrum and a balanced investigation of partial and full delegation.

  • Its stated aim is a balanced scientific discussion of using LLMs for relevance judgments rather than a one-sided position.
  • Cranfield-style evaluation requires human relevance judgments, making test-collection construction time-intensive and costly.
  • The paper asks whether LLMs can partially or fully delegate relevance judgment while acknowledging uncertainty about agreement with human annotators.
  • The authors propose a spectrum from manual judgments to fully automated judgments, with varying levels of human involvement and decision making.
  • The paper discusses existing and emerging human–machine collaboration scenarios, associated risks, open questions, and a pilot agreement experiment.

2 RELATED WORK

Related work shows that IR has developed multiple ways to reduce or redistribute the cost of relevance assessment. These approaches vary from supporting human assessors to inferring or replacing document-level judgments.

  • Assessment Systems: Human-assessment tools support annotation through features such as text highlighting, pre-annotations, and external data integration.
  • Crowdsourcing: Crowdsourcing emerged as collections grew, prompting research into the reliability, cost, and quality management of crowdsourced judgments.
  • Some methods retain human control over relevance while allowing machines to select documents or derive assessments.
  • Passage-ROUGE and BertScore: Similarity-based approaches can transfer labels to unjudged passages, using measures such as ROUGE or BertScore.
  • Other strategies reduce assessment needs by adjusting metrics, using exam-question answerability, predicting query performance, or generating queries automatically.
  • Reconstruct Documents: Document reconstruction derives relevance from human-authored structure such as anchor text, metadata, categories, glosses, or infoboxes.
  • Evaluation of Automatic Evaluation: Prior evaluation work examined agreement between automatic and manual assessments through leaderboard correlation, motivating the paper’s analogous LLM study.

3 SPECTRUM OF HUMAN–MACHINE COLLABORATION

The paper organizes relevance-judgment strategies along a four-level human–machine collaboration spectrum, from manual assessment to full LLM automation. Intermediate levels assign different judgment subtasks to humans and machines, including assistance, verification, and competence partitioning.

  • The spectrum ranges from humans making judgments manually to LLMs replacing humans completely, with intermediate collaboration levels.
  • Human Judgment: Human Judgment keeps relevance decisions with people while allowing basic interface support such as highlighting or document clustering.
  • Human Verification: Two LLMs can independently generate judgments while humans select the better one, preserving human involvement in machine-supported assessment.
  • AI Assistance: AI Assistance lets humans judge documents using LLM-generated summaries or automated counts of manually defined information nuggets.
  • The spectrum raises research questions about which subtasks need human input, how LLM assistance affects assessors, and whether humans can eventually be replaced.
  • Human Verification: Human Verification uses an LLM’s first-pass judgment and rationale as a suggestion that a human accepts, rejects, or reviews selectively.
  • Fully Automated: Fully Automated judgments would replace humans if LLMs reliably produced high-quality relevance assessments for a specified corpus or domain.
  • The paper proposes competence partitioning, assigning subtasks to humans or machines according to which performs them better, with humans currently viewed as the only reliable final judges.

4 OPEN ISSUES, FORESEEABLE RISKS, AND OPPORTUNITIES

The paper maps open issues, risks, and opportunities for using LLMs in relevance judgments, from human assistance to fully automated evaluation. It highlights uncertain judgment quality, bias, truthfulness, circularity, and limits of current evaluation assumptions.

  • LLMs’ Judgment Quality: LLM-based relevance judgments may greatly increase annotation volume while reducing quality, although the extent of deterioration remains unclear.The paper calls for LLM-specific quality assurance and considers active learning in which models learn from human decisions.
  • Human–Machine Collaboration: The human–machine spectrum distinguishes assessor types and judgment tasks to identify where LLMs might replace or support users, experts, crowdworkers, or other LLMs.Tasks include preference, binary, graded, and explained relevance judgments.
  • Using Multiple LLMs as Assessors: Multiple LLM assessors may produce correlated answers because similar training corpora can yield agreement whose correctness is unknown.Using distinct subcorpora and personalized models is proposed as one possible response.
  • Truthfulness & Misinformation: Fully automatic assessment must address factuality because topical relevance can diverge from whether a document correctly answers the information need.The setting would require LLMs to verify sources and document truthfulness, raising unresolved fact-checking questions.
  • Bias and Explaining Relevance: LLM bias may enter relevance judgments, while current systems may struggle with style, truthfulness, cultural expectations, and other task-specific aspects of relevance.The paper therefore identifies human intervention as necessary for collecting and judging additional facts or document aspects not easily discerned by LLMs.
  • Moving Beyond Cranfield and Human: Using the same LLM for retrieval and evaluation could create circularity, while moving beyond human assessors could make human-annotated gold standards unable to detect superior LLM performance.The paper also questions assumptions such as static collections and judgments that do not depend on other ranked documents.

5 PRELIMINARY ASSESSMENT

The pilot compares human and LLM relevance judgments across two TREC collections, judgment types, prompts, and models. Agreement varies substantially by collection and relevance label, with LLMs aligning better for highly relevant graded passages than for binary relevant documents.

  • Experimental Setup: The pilot compares GPT-3.5 and YouChat with TREC assessors on TREC-8 and TREC-DL 2021 using binary and graded judgments.The experiments used tailored prompts and were conducted in January and February 2023.
  • Experimental Setup: The experiments are exploratory rather than exhaustive, aiming to identify where LLM judgments agree or disagree with manual relevance judgments.The authors frame the study as a preliminary assessment of current LLM capability.
  • TREC-8 Results: 90% of human non-relevant TREC-8 documents received the same GPT-3.5 label, compared with 50% of human-relevant documents.YouChat agreement was 74% for non-relevant documents and 33% for relevant documents.
  • Interpretation: The authors hypothesize that humans better recognize subtle relevance distinctions, whereas LLMs may correlate better with coarse-grained graded judgments.They suggest centering graded judgments symmetrically around 0 as a possible direction.

6 RE-JUDGING TREC 2021 DEEP LEARNING

The re-judging experiment applies GPT-3.5 judgments to TREC-DL 2021 passage-ranking runs and compares resulting system rankings and discriminative power with official human judgments. The top-ranked system is preserved, but LLM-based measures distinguish fewer run pairs.

  • Experimental Setup: The experiment re-judges the TREC-DL 2021 passage-ranking pool with GPT-3.5 using graded judgments to approximate the track’s methodology.It complements the earlier binary-focused experiments.
  • Experimental Setup: The GPT-3.5 prompt uses few-shot examples for multiple relevance levels, and the re-judgments cost around USD 1 cent each.The total reported expenditure was USD 111.90, including duplicate requests.
  • Evaluation: Kendall’s τ measures the correlation between system rankings computed from official human judgments and GPT-3.5 judgments.The analysis reports both the full four-point scale and the binary convention used for MAP.
  • Results: The top run under official judgments remains the top run under LLM judgments.The paper compares this result with a reported Kendall’s τ = .90 for MAP from a similar experiment using two human judgment types.
  • Results: 72% of run pairs are distinguished under human MAP judgments versus 65% under GPT-3.5 judgments; for NDCG@10, the figures are 74% and 69%.Sensitivity is defined using paired t-tests with p < 0.05, without correcting for multiple comparisons.

7 PERSPECTIVES FOR THE FUTURE

The paper presents perspectives for, against, and between fully automated LLM relevance judgments. It finds promise in scalable assistance and double-checking, but emphasizes unresolved concerns about human grounding, circularity, bias, and credibility.

  • Perspectives: The paper frames its future discussion as opposing and compromise perspectives informed by literature analysis, pilot evidence, and IR research experience.The perspectives address both automatic judgment and human–machine collaboration.
  • In Favor of Using LLMs: LLMs can generate relevance explanations that assist human assessors, while human decisions can provide quality control and feedback for model improvement.The paper warns that generated labels and explanations may also bias or mislead assessors.
  • In Favor of Using LLMs: The paper identifies scalability, consistency, multilingual processing, multimodal assessment, and explanations as potential advantages of LLM assessors.These capabilities could support larger and more complex test collections.
  • Against Using LLMs: Fully automated LLM judgments lack proof of a relationship to real-world user relevance, which is subjective and can change over time.The paper therefore treats replacement of human assessors as requiring further research before deployment.
  • Against Using LLMs: Using an LLM for both retrieval and evaluation could inflate measured performance and penalize systems based on different relevance rationales.This is the paper’s central circularity concern for LLM-based evaluation.
  • A Compromise: The paper favors human–machine collaboration, especially AI assistance and human verification, while noting that these approaches remain comparatively underexplored.Automatic judgments may still help evaluate early prototypes, initiate novel-task judgments, and support large-scale training.

8 CONCLUSION

The paper finds that LLMs show promise for mimicking human relevance assessments and offers a starting point for further research on their role in IR evaluation.

  • LLMs show promise in mimicking human relevance assessments.
  • The paper provides a starting point for future research on LLMs for relevance judgment.
  • The authors present views both for and against employing LLMs in the IR evaluation process.
Loading 2304.09161v2…