paper-with-me

홈 › Papers

Metric-Dependent Annotation Saturation for Learning from Label Distributions

2026-05-28 · Guneet Kohli arxiv

When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distributions subsampled from ChaosNLI, a dataset providing 100 independent annotator judgments per item, and identify metric-dependent saturation. In our 3-class NLI setting, entropy correlation -- whether the model identifies which items elicit disagreement -- requires N ~ 20-50 annotators to converge, while distributional match (KL divergence) saturates by N ~ 10 (87-95% of improvement across five model seeds). This finding rests on a prior observation: soft labels carry item-specific signal that label smoothing cannot replicate. Across five smoothing intensities, entropy correlation clusters at r ~ 0.45-0.49, while soft labels reach r = 0.643 (p < 0.001); per-item analysis traces this gap to smoothing's inability to distinguish ambiguous items from clear ones. The soft-label advantage replicates across two architectures (DeBERTa, RoBERTa), a non-NLI-pretrained baseline, and an exploratory cross-domain evaluation on content safety. These results suggest that annotation budgets should be informed by the target evaluation metric rather than set uniformly.

📄 PDF Abstract BibTeX arXiv:2605.29797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Clinical Uncertainty Impacts Machine Learning Evaluations

2025-09-26 · Simone Lionetti, Fabian Gröger, Philippe Gottfrois, Alvaro Gonzalez-Jimenez 외 arxiv

Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on me…

Exploring the Properties and Evolution of Neural Network Eigenspaces during Training

2021-06-17 · Mats L. Richter, Leila Malihi, Anne-Kathrin Patricia Windler, Ulf Krumnack

In this work we explore the information processing inside neural networks using logistic regression probes \cite{probes} and the saturation metric \cite{featurespace_saturation}. We show that problem difficulty and neura…

regression

The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR

2025-10-26 · Siyu Liang, Nicolas Ballier, Gina-Anne Levow, Richard Wright arxiv

How much audio is needed to fully observe a multilingual ASR model's learned sub-token inventory across languages, and does data disparity in multilingual pre-training affect how these tokens are utilized during inferenc…

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

2026-07-19 · Pilsung Kang arxiv

Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop …

Color image restoration based on nonlocal saturation-value similarity

2026-03-19 · Wei Wang, Yakun Li arxiv

In this paper, we propose and develop a novel nonlocal variational technique based on saturation-value similarity for color image restoration. In traditional nonlocal methods, image patches are extracted from red, green …

Image Restoration