paper-with-me

홈 › Papers

Redefining Retrieval Evaluation in the Era of LLMs

2025-10-24 · Giovanni Trappolini, Florin Cuconasu, Simone Filice, Yoelle Maarek, Fabrizio Silvestri arxiv

Traditional Information Retrieval (IR) metrics, such as nDCG, MAP, and MRR, assume that human users sequentially examine documents with diminishing attention to lower ranks. This assumption breaks down in Retrieval Augmented Generation (RAG) systems, where search results are consumed by Large Language Models (LLMs), which, unlike humans, process all retrieved documents as a whole rather than sequentially. Additionally, traditional IR metrics do not account for related but irrelevant documents that actively degrade generation quality, rather than merely being ignored. Due to these two major misalignments, namely human vs. machine position discount and human relevance vs. machine utility, classical IR metrics do not accurately predict RAG performance. We introduce a utility-based annotation schema that quantifies both the positive contribution of relevant passages and the negative impact of distracting ones. Building on this foundation, we propose UDCG (Utility and Distraction-aware Cumulative Gain), a metric using an LLM-oriented positional discount to directly optimize the correlation with the end-to-end answer accuracy. Experiments on five datasets and six LLMs demonstrate that UDCG improves correlation by up to 36% compared to traditional metrics. Our work provides a critical step toward aligning IR evaluation with LLM consumers and enables more reliable assessment of RAG components

📄 PDF Abstract BibTeX arXiv:2510.21440

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies

2025-03-10 · Luyi Jiang, Jiayuan Chen, Lu Lu, Xinwei Peng 외

The evaluation and improvement of medical large language models (LLMs) are critical for their real-world deployment, particularly in ensuring accuracy, safety, and ethical alignment. Existing frameworks inadequately diss…

BenchmarkingEthicsHallucinationPrompt Engineering+1

GOLF: Goal-Oriented Long-term liFe tasks supported by human-AI collaboration

2024-03-25 · Ben Wang

The advent of ChatGPT and similar large language models (LLMs) has revolutionized the human-AI interaction and information-seeking process. Leveraging LLMs as an alternative to search engines, users can now access summar…

Decision MakingInformation RetrievalTask Planning

Redefining Information Retrieval of Structured Database via Large Language Models

2024-05-09 · Mingzhu Wang, Yuzhe Zhang, Qihang Zhao, Junyi Yang 외

Retrieval augmentation is critical when Language Models (LMs) exploit non-parametric knowledge related to the query through external knowledge bases before reasoning. The retrieved information is incorporated into LMs as…

Information RetrievalQuestion AnsweringRetrieval

A Survey of LLM $\times$ DATA

2025-05-24 · Xuanhe Zhou, Junxuan He, Wei Zhou, Haodong Chen 외

The integration of large language model (LLM) and data management (DATA) is rapidly redefining both domains. In this survey, we comprehensively review the bidirectional relationships. On the one hand, DATA4LLM, spanning …

Large Language ModelManagementRAGRetrieval+2

Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation

2026-04-16 · Bhavik Vachhani, Kush Shrisvastava, Pranshu Nema, Sai Chiranthan arxiv

Evaluating large language models (LLMs) for clinical documentation tasks such as SOAP note generation remains challenging. Unlike standard summarization, these tasks require clinical abstraction, normalization of colloqu…