paper-with-me

Papers

Using Score Distributions to Compare Statistical Significance Tests for Information Retrieval Evaluation

2019-01-30 · Parapar Javier, Losada David E., Presedo-Quindimil Manuel A., Barreiro Alvaro

Statistical significance tests can provide evidence that the observed difference in performance between two methods is not due to chance. In Information Retrieval, some studies have examined the validity and suitability of such tests for comparing search systems. We argue here that current methods for assessing the reliability of statistical tests suffer from some methodological weaknesses, and we propose a novel way to study significance tests for retrieval evaluation. Using Score Distributions, we model the output of multiple search systems, produce simulated search results from such models, and compare them using various significance tests. A key strength of this approach is that we assess statistical tests under perfect knowledge about the truth or falseness of the null hypothesis. This new method for studying the power of significance tests in Information Retrieval evaluation is formal and innovative. Following this type of analysis, we found that both the sign test and Wilcoxon signed test have more power than the permutation test and the t-test. The sign test and Wilcoxon signed test also have a good behavior in terms of type I errors. The bootstrap test shows few type I errors, but it has less power than the other methods tested.

📄 PDF Abstract BibTeX arXiv:1901.10696

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

NLPStatTest: A Toolkit for Comparing NLP System Performance

2020-11-26 · Asian Chapter of the Association for Computational Linguistics 2020 · Haotian Zhu, Denise Mak, Jesse Gioannini, Fei Xia

Statistical significance testing centered on p-values is commonly used to compare NLP system performance, but p-values alone are insufficient because statistical significance differs from practical significance. The latt…

A Hitchhiker's Guide to Statistical Comparisons of Reinforcement Learning Algorithms

2019-04-15 · Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer

Consistently checking the statistical significance of experimental results is the first mandatory step towards reproducible science. This paper presents a hitchhiker's guide to rigorous comparisons of reinforcement learn…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

A Hitchhiker's Guide to Statistical Comparisons of Reinforcement Learning Algorithms

2019-03-06 · ICLR Workshop RML 2019 5 · Anonymous

Consistently checking the statistical significance of experimental results is the first mandatory step towards reproducible science. This paper presents a hitchhiker's guide to rigorous comparisons of reinforcement learn…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Unveiling Statistical Significance of Online Regression over Multiple Datasets

2025-12-14 · Mohammad Abu-Shaira, Weishi Shi arxiv

Despite extensive focus on techniques for evaluating the performance of two learning algorithms on a single dataset, the critical challenge of developing statistical tests to compare multiple algorithms across various da…

Significance Tests for Neural Networks

2019-02-16 · Enguerrand Horel, Kay Giesecke

We develop a pivotal test to assess the statistical significance of the feature variables in a single-layer feedforward neural network regression model. We propose a gradient-based test statistic and study its asymptotic…

Computational Efficiencyregression