paper-with-me

홈 › Papers

When Experts Disagree: Characterizing Annotator Variability for Vessel Segmentation in DSA Images

2025-08-14 · M. Geshvadi, G. So, D. D. Chlorogiannis, C. Galvin, E. Torio, A. Azimi, Y. Tachie-Baffour, N. Haouchine, A. Golby, M. Vangel, W. M. Wells, Y. Epelboym, R. Du, F. Durupinar, S. Frisken arxiv

We analyze the variability among segmentations of cranial blood vessels in 2D DSA performed by multiple annotators in order to characterize and quantify segmentation uncertainty. We use this analysis to quantify segmentation uncertainty and discuss ways it can be used to guide additional annotations and to develop uncertainty-aware automatic segmentation methods.

📄 PDF Abstract BibTeX arXiv:2508.10797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NUTMEG: Separating Signal From Noise in Annotator Disagreement

2025-07-25 · Jonathan Ivey, Susan Gauch, David Jurgens arxiv

NLP models often rely on human-labeled data for training and evaluation. Many approaches crowdsource this data from a large number of annotators with varying skills, backgrounds, and motivations, resulting in conflicting…

Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the Truth

2020-05-01 · ICLR 2020 1 · Igor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei 외

In most machine learning tasks unambiguous ground truth labels can easily be acquired. However, this luxury is often not afforded to many high-stakes, real-world scenarios such as medical image interpretation, where even…

BIG-bench Machine LearningBinary Classification

Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives

2025-08-04 · Yinuo Xu, Veronica Derricks, Allison Earl, David Jurgens arxiv

We present an approach to modeling annotator disagreement in subjective NLP tasks through both architectural and data-centric innovations. Our model, DEM-MoE (Demographic-Aware Mixture of Experts), routes inputs to exper…

Learning Ambiguity from Crowd Sequential Annotations

2023-01-04 · Xiaolei Lu

Most crowdsourcing learning methods treat disagreement between annotators as noisy labelings while inter-disagreement among experts is often a good indicator for the ambiguity and uncertainty that is inherent in natural …

NERPOSPOS Tagging

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

2026-05-04 · Akash Bonagiri, Gerard Janno Anderias, Saee Patil, Angelina Lai 외 arxiv

Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discar…