paper-with-me

홈 › Papers

Relevance Assessments for Web Search Evaluation: Should We Randomise or Prioritise the Pooled Documents? (CORRECTED VERSION)

2022-11-02 · Tetsuya Sakai, Sijie Tao, Zhaohao Zeng

In the context of depth-$k$ pooling for constructing web search test collections, we compare two approaches to ordering pooled documents for relevance assessors: the prioritisation strategy (PRI) used widely at NTCIR, and the simple randomisation strategy (RND). In order to address research questions regarding PRI and RND, we have constructed and released the WWW3E8 data set, which contains eight independent relevance labels for 32,375 topic-document pairs, i.e., a total of 259,000 labels. Four of the eight relevance labels were obtained from PRI-based pools; the other four were obtained from RND-based pools. Using WWW3E8, we compare PRI and RND in terms of inter-assessor agreement, system ranking agreement, and robustness to new systems that did not contribute to the pools. We also utilise an assessor activity log we obtained as a byproduct of WWW3E8 to compare the two strategies in terms of assessment efficiency.

📄 PDF Abstract BibTeX arXiv:2211.00981

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

A Blueprint of IR Evaluation Integrating Task and User Characteristics: Test Collection and Evaluation Metrics

2023-05-01 · Kal Jarvelin, Eero Sormunen

Relevance is generally understood as a multi-level and multi-dimensional relationship between an information need and an information object. However, traditional IR evaluation metrics naively assume mono-dimensionality. …

LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?

2024-11-11 · Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian Soboroff

Test collections are information retrieval tools that allow researchers to quickly and easily evaluate ranking algorithms. While test collections have become an integral part of IR research, the process of data creation …

Information Retrieval

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation

2024-11-20 · David Otero, Javier Parapar, Álvaro Barreiro

Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating which documents are relevant for each topi…

Information RetrievalRetrieval

Augmented Test Collections: A Step in the Right Direction

2015-01-26 · Hasler Laura, Halvey Martin, Villa Robert

In this position paper we argue that certain aspects of relevance assessment in the evaluation of IR systems are oversimplified and that human assessments represented by qrels should be augmented to take account of conte…

Position

One-Shot Labeling for Automatic Relevance Estimation

2023-02-22 · Sean MacAvaney, Luca Soldaini

Dealing with unjudged documents ("holes") in relevance assessments is a perennial problem when evaluating search systems with offline experiments. Holes can reduce the apparent effectiveness of retrieval systems during e…

Retrieval