paper-with-me

Papers

Perspectives on Large Language Models for Relevance Judgment

2023-04-13 · Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, Henning Wachsmuth

When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems. In this perspectives paper, we discuss possible ways for LLMs to support relevance judgments along with concerns and issues that arise. We devise a human--machine collaboration spectrum that allows to categorize different relevance judgment strategies, based on how much humans rely on machines. For the extreme point of "fully automated judgments", we further include a pilot experiment on whether LLM-based relevance judgments correlate with judgments from trained human assessors. We conclude the paper by providing opposing perspectives for and against the use of~LLMs for automatic relevance judgments, and a compromise perspective, informed by our analyses of the literature, our preliminary experimental evidence, and our experience as IR researchers.

📄 PDF Abstract BibTeX arXiv:2304.09161

Code (1)

narabzad/llm-relevance-judgement-comparison

Tasks

Retrieval

Similar Papers 제목 키워드 기반

LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations

2025-04-27 · Laura Dietz, Oleg Zendel, Peter Bailey, Charles Clarke 외

Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based eva…

Information Retrieval

Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval

2024-03-27 · Shengjie Ma, Chong Chen, Qi Chu, Jiaxin Mao

Collecting relevant judgments for legal case retrieval is a challenging and time-consuming task. Accurately judging the relevance between two legal cases requires a considerable effort to read the lengthy text and a high…

Language ModelingLanguage ModellingLarge Language ModelRetrieval

Iterative Utility Judgment Framework via LLMs Inspired by Relevance in Philosophy

2024-06-17 · Hengran Zhang, Keping Bi, Jiafeng Guo, Xueqi Cheng

Utility and topical relevance are critical measures in information retrieval (IR), reflecting system and user perspectives, respectively. While topical relevance has long been emphasized, utility is a higher standard of …

Answer GenerationInformation RetrievalPassage RetrievalPhilosophy+4

How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective

2025-04-10 · Qi Liu, Jiaxin Mao, Ji-Rong Wen

Recent studies have shown that large language models (LLMs) can assess relevance and support information retrieval (IR) tasks such as document ranking and relevance judgment generation. However, the internal mechanisms b…

Document RankingInformation Retrieval

Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation

2024-08-02 · Jheng-Hong Yang, Jimmy Lin

Vision--Language Models (VLMs) have demonstrated success across diverse applications, yet their potential to assist in relevance judgments remains uncertain. This paper assesses the relevance estimation capabilities of V…

Image-text RetrievalRetrievalText Retrieval