paper-with-me

홈 › Papers

On Survivorship Bias in MS MARCO

2022-04-27 · Prashansa Gupta, Sean MacAvaney

Survivorship bias is the tendency to concentrate on the positive outcomes of a selection process and overlook the results that generate negative outcomes. We observe that this bias could be present in the popular MS MARCO dataset, given that annotators could not find answers to 38--45% of the queries, leading to these queries being discarded in training and evaluation processes. Although we find that some discarded queries in MS MARCO are ill-defined or otherwise unanswerable, many are valid questions that could be answered had the collection been annotated more completely (around two thirds using modern ranking techniques). This survivability problem distorts the MS MARCO collection in several ways. We find that it affects the natural distribution of queries in terms of the type of information needed. When used for evaluation, we find that the bias likely yields a significant distortion of the absolute performance scores observed. Finally, given that MS MARCO is frequently used for model training, we train models based on subsets of MS MARCO that simulates more survivorship bias. We find that models trained in this setting are up to 9.9% worse when evaluated on versions of the dataset with more complete annotations, and up to 3.5% worse at zero-shot transfer. Our findings are complementary to other recent suggestions for further annotation of MS MARCO, but with a focus on discarded queries. Code and data for reproducing the results of this paper are available in an online appendix.

📄 PDF Abstract BibTeX arXiv:2204.12852

Code (1)

seanmacavaney/sigir2022-survivor 공식 구현

Tasks

valid

Similar Papers 제목 키워드 기반

Better Protein Function Prediction by Modeling Survivorship Bias

2026-05-07 · Zhongmou Chao, Poompol Buathong, Ekaterina Selivanovitch, Susan Daniel 외 arxiv

Protein sequence data from nature exhibits survivorship bias: we only observe data from those organisms that survive and reproduce, while non-functional protein mutations are eliminated by natural selection. Thus, predic…

Protein Function Prediction

Lineage EM Algorithm for Inferring Latent States from Cellular Lineage Trees

2019-11-29

Phenotypic variability in a population of cells can work as the bet-hedging of the cells under an unpredictably changing environment, the typical example of which is the bacterial persistence. To understand the strategy …

Understanding Performance of Long-Document Ranking Models through Comprehensive Evaluation and Leaderboarding

2022-07-04 · Leonid Boytsov, David Akinpelu, Tianyi Lin, Fangwei Gao 외

We evaluated 20+ Transformer models for ranking of long documents (including recent LongP models trained with FlashAttention) and compared them with a simple FirstP baseline, which applies the same model to the truncated…

BenchmarkingDocument Ranking

Evaluating LLMs in Finance Requires Explicit Bias Consideration

2026-02-15 · Yaxuan Kong, Hoyoung Lee, Yoontae Hwang, Alejandro Lopez-Lira 외 arxiv

Large Language Models (LLMs) are increasingly integrated into financial workflows, but evaluation practice has not kept up. Finance-specific biases can inflate performance, contaminate backtests, and make reported result…

Log2NS: Enhancing Deep Learning Based Analysis of Logs With Formal to Prevent Survivorship Bias

2021-05-29 · Charanraj Thimmisetty, Praveen Tiwari, Didac Gil de la Iglesia, Nandini Ramanan 외

Analysis of large observational data sets generated by a reactive system is a common challenge in debugging system failures and determining their root cause. One of the major problems is that these observational data suf…