Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets
With the ever-growing amounts of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in order to ensure consistent performance across heterogeneous setups. However, such multiple comparisons pose significant challenges to traditional statistical analysis methods in NLP and can lead to erroneous conclusions. In this paper, we propose a Replicability Analysis framework for a statistically sound analysis of multiple comparisons between algorithms for NLP tasks. We discuss the theoretical advantages of this framework over the current, statistically unjustified, practice in the NLP literature, and demonstrate its empirical value across four applications: multi-domain dependency parsing, multilingual POS tagging, cross-domain sentiment classification and word similarity prediction.
Code (1)
Tasks
Dependency ParsingGeneral ClassificationPOSPOS TaggingSentiment AnalysisSentiment ClassificationWord SimilaritySimilar Papers 제목 키워드 기반
Replicability of Research in Biomedical Natural Language Processing: a pilot evaluation for a coding task
Community Perspective on Replicability in Natural Language Processing
With recent efforts in drawing attention to the task of replicating and/or reproducing results, for example in the context of COLING 2018 and various LREC workshops, the question arises how the NLP community views the to…
SurveyNondistributivity of human logic and violation of response replicability effect in cognitive psychology
The aim of this paper is to promote quantum logic as one of the basic tools for analyzing human reasoning. We compare it with classical (Boolean) logic and highlight the role of violation of the distributive law for conj…
Replicability in High Dimensional Statistics
The replicability crisis is a major issue across nearly all areas of empirical science, calling for the formal study of replicability in statistics. Motivated in this context, [Impagliazzo, Lei, Pitassi, and Sorrell STOC…
Replicable Uniformity Testing
Uniformity testing is arguably one of the most fundamental distribution testing problems. Given sample access to an unknown distribution $\mathbf{p}$ on $[n]$, one must decide if $\mathbf{p}$ is uniform or $\varepsilon$-…