paper-with-me

홈 › Papers

One-Shot Labeling for Automatic Relevance Estimation

2023-02-22 · Sean MacAvaney, Luca Soldaini

Dealing with unjudged documents ("holes") in relevance assessments is a perennial problem when evaluating search systems with offline experiments. Holes can reduce the apparent effectiveness of retrieval systems during evaluation and introduce biases in models trained with incomplete data. In this work, we explore whether large language models can help us fill such holes to improve offline evaluations. We examine an extreme, albeit common, evaluation setting wherein only a single known relevant document per query is available for evaluation. We then explore various approaches for predicting the relevance of unjudged documents with respect to a query and the known relevant document, including nearest neighbor, supervised, and prompting techniques. We find that although the predictions of these One-Shot Labelers (1SL) frequently disagree with human assessments, the labels they produce yield a far more reliable ranking of systems than the single labels do alone. Specifically, the strongest approaches can consistently reach system ranking correlations of over 0.86 with the full rankings over a variety of measures. Meanwhile, the approach substantially increases the reliability of t-tests due to filling holes in relevance assessments, giving researchers more confidence in results they find to be significant. Alongside this work, we release an easy-to-use software package to enable the use of 1SL for evaluation of other ad-hoc collections or systems.

📄 PDF Abstract BibTeX arXiv:2302.11266

Code (1)

seanmacavaney/autoqrels 공식 구현

Tasks

Retrieval

Similar Papers 제목 키워드 기반

MetricPrompt: Prompting Model as a Relevance Metric for Few-shot Text Classification

2023-06-15 · Hongyuan Dong, Weinan Zhang, Wanxiang Che

Prompting methods have shown impressive performance in a variety of text mining tasks and applications, especially few-shot ones. Despite the promising prospects, the performance of prompting model largely depends on the…

ClassificationFew-Shot Text Classificationtext-classificationText Classification

Accuracy of Automatic Cross-Corpus Emotion Labeling for Conversational Speech Corpus Commonization

2016-05-01 · LREC 2016 5 · Hiroki Mori, Atsushi Nagaoka, Yoshiko Arimoto

There exists a major incompatibility in emotion labeling framework among emotional speech corpora, that is, category-based and dimension-based. Commonizing these requires inter-corpus emotion labeling according to both f…

Cross-corpus

Learning text-to-video retrieval from image captioning

2024-04-26 · Lucas Ventura, Cordelia Schmid, Gül Varol

We describe a protocol to study text-to-video retrieval training with unlabeled videos, where we assume (i) no access to labels for any videos, i.e., no access to the set of ground-truth captions, but (ii) access to labe…

Image CaptioningImage RetrievalRetrievalText to Video Retrieval+2

Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation

2024-08-02 · Jheng-Hong Yang, Jimmy Lin

Vision--Language Models (VLMs) have demonstrated success across diverse applications, yet their potential to assist in relevance judgments remains uncertain. This paper assesses the relevance estimation capabilities of V…

Image-text RetrievalRetrievalText Retrieval

Domain Adaptation for Dense Retrieval and Conversational Dense Retrieval through Self-Supervision by Meticulous Pseudo-Relevance Labeling

2024-03-13 · Minghan Li, Eric Gaussier

Recent studies have demonstrated that the ability of dense retrieval models to generalize to target domains with different distributions is limited, which contrasts with the results obtained with interaction-based models…

Conversational SearchDomain AdaptationRetrieval