LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation
LongEval-Retrieval is a Web document retrieval benchmark that focuses on continuous retrieval evaluation. This test collection is intended to be used to study the temporal persistence of Information Retrieval systems and will be used as the test collection in the Longitudinal Evaluation of Model Performance Track (LongEval) at CLEF 2023. This benchmark simulates an evolving information system environment - such as the one a Web search engine operates in - where the document collection, the query distribution, and relevance all move continuously, while following the Cranfield paradigm for offline evaluation. To do that, we introduce the concept of a dynamic test collection that is composed of successive sub-collections each representing the state of an information system at a given time step. In LongEval-Retrieval, each sub-collection contains a set of queries, documents, and soft relevance assessments built from click models. The data comes from Qwant, a privacy-preserving Web search engine that primarily focuses on the French market. LongEval-Retrieval also provides a 'mirror' collection: it is initially constructed in the French language to benefit from the majority of Qwant's traffic, before being translated to English. This paper presents the creation process of LongEval-Retrieval and provides baseline runs and analysis.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalPrivacy PreservingRetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
This paper presents the third edition of the LongEval Lab, part of the CLEF 2025 conference, which continues to explore the challenges of temporal persistence in Information Retrieval (IR). The lab features two tasks des…
Information RetrievalRetrievalCross-Lingual Training with Dense Retrieval for Document Retrieval
Dense retrieval has shown great success in passage ranking in English. However, its effectiveness in document retrieval for non-English languages remains unexplored due to the limitation in training resources. In this wo…
Document RankingPassage RankingRetrievalSubmitted and Diagnostic Analysis of Full-Text Temporal Retrieval for LongEval-Sci
LongEval-Sci evaluates scientific retrieval under collection change, where a system should be effective on the current corpus and remain usable as documents accumulate over time. This paper reports both official Task 1 r…
Text RetrievalCURE: A dataset for Clinical Understanding & Retrieval Evaluation
Given the dominance of dense retrievers that do not generalize well beyond their training dataset distributions, domain-specific test sets are essential in evaluating retrieval. There are few test datasets for retrieval …
Passage RankingRetrievalAnalyzing the Effectiveness of Listwise Reranking with Positional Invariance on Temporal Generalizability
This working note outlines our participation in the retrieval task at CLEF 2024. We highlight the considerable gap between studying retrieval performance on static knowledge documents and understanding performance in rea…
BenchmarkingDecoderInformation RetrievalReranking+1