paper-with-me

홈 › Papers

We Need to Talk about Standard Splits

2019-07-01 · ACL 2019 7 · Kyle Gorman, Steven Bedrick

It is standard practice in speech {\&} language technology to rank systems according to their performance on a test set held out for evaluation. However, few researchers apply statistical tests to determine whether differences in performance are likely to arise by chance, and few examine the stability of system ranking across multiple training-testing splits. We conduct replication and reproduction experiments with nine part-of-speech taggers published between 2000 and 2018, each of which claimed state-of-the-art performance on a widely-used {``}standard split{''}. While we replicate results on the standard split, we fail to reliably reproduce some rankings when we repeat this analysis with randomly generated training-testing splits. We argue that randomly generated splits should be used in system evaluation.

📄 PDF Abstract BibTeX

Code (1)

kylebgorman/SOTA-taggers 공식 구현

Similar Papers 제목 키워드 기반

We Need to Talk About Random Splits

2020-05-01 · EACL 2021 2 · Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, Katja Filippova

Gorman and Bedrick (2019) argued for using random splits rather than standard splits in NLP experiments. We argue that random splits, like standard splits, lead to overly optimistic performance estimates. We can also spl…

Domain Adaptation

We Need to Talk About train-dev-test Splits

2021-11-01 · EMNLP 2021 11 · Rob van der Goot

Standard train-dev-test splits used to benchmark multiple models against each other are ubiquitously used in Natural Language Processing (NLP). In this setup, the train data is used for training the model, the developmen…

Model Selection

Auditory distraction in open-plan office environments: The effect of multi-talker acoustics

2023-04-14 · Manuj Yadav, Jungsoo Kim, Densil Cabrera, Richard de Dear

Within the soundscapes of open-plan offices, irrelevant speech has consistently been reported as the most distracting, and causing performance decrements for workers. Notwithstanding this generalization, the 'babble' cre…

Experimental Design

K-Splits: Improved K-Means Clustering Algorithm to Automatically Detect the Number of Clusters

2021-10-09 · Seyed Omid Mohammadi, Ahmad Kalhor, Hossein Bodaghi

This paper introduces k-splits, an improved hierarchical algorithm based on k-means to cluster data without prior knowledge of the number of clusters. K-splits starts from a small number of clusters and uses the most sig…

ClusteringPosition

TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection

2025-05-30 · Xinqi Xiong, Prakrut Patel, Qingyuan Fan, Amisha Wadhwa 외

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and …

DeepFake DetectionFace SwappingHead Detection