On the Reliability of Test Collections for Evaluating Systems of Different Types
As deep learning based models are increasingly being used for information retrieval (IR), a major challenge is to ensure the availability of test collections for measuring their quality. Test collections are generated based on pooling results of various retrieval systems, but until recently this did not include deep learning systems. This raises a major challenge for reusable evaluation: Since deep learning based models use external resources (e.g. word embeddings) and advanced representations as opposed to traditional methods that are mainly based on lexical similarity, they may return different types of relevant document that were not identified in the original pooling. If so, test collections constructed using traditional methods are likely to lead to biased and unfair evaluation results for deep learning (neural) systems. This paper uses simulated pooling to test the fairness and reusability of test collections, showing that pooling based on traditional systems only can lead to biased evaluation of deep learning systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningFairnessInformation RetrievalRetrievalWord EmbeddingsSimilar Papers 제목 키워드 기반
Intelligent Topic Selection for Low-Cost Information Retrieval Evaluation: A New Perspective on Deep vs. Shallow Judging
While test collections provide the cornerstone for Cranfield-based evaluation of information retrieval (IR) systems, it has become practically infeasible to rely on traditional pooling techniques to construct test collec…
Information RetrievalRetrievalTowards Understanding Bias in Synthetic Data for Evaluation
Test collections are crucial for evaluating Information Retrieval (IR) systems. Creating a diverse set of user queries for these collections can be challenging, and obtaining relevance judgments, which indicate how well …
Information RetrievalREANIMATOR: Reanimate Retrieval Test Collections with Extracted and Synthetic Resources
Retrieval test collections are essential for evaluating information retrieval systems, yet they often lack generalizability across tasks. To overcome this limitation, we introduce REANIMATOR, a versatile framework design…
Information RetrievalRetrievalRetrieval-augmented GenerationEvaluating Machine Reading Systems through Comprehension Tests
This paper describes a methodology for testing and evaluating the performance of Machine Reading systems through Question Answering and Reading Comprehension Tests. The methodology is being used in QA4MRE (QA for Machine…
Answer SelectionMultiple-choiceQuestion AnsweringReading ComprehensionGenTREC: The First Test Collection Generated by Large Language Models for Evaluating Information Retrieval Systems
Building test collections for Information Retrieval evaluation has traditionally been a resource-intensive and time-consuming task, primarily due to the dependence on manual relevance judgments. While various cost-effect…
Information RetrievalLarge Language ModelRetrieval