Variability Matters : Evaluating inter-rater variability in histopathology for robust cell detection
Large annotated datasets have been a key component in the success of deep learning. However, annotating medical images is challenging as it requires expertise and a large budget. In particular, annotating different types of cells in histopathology suffer from high inter- and intra-rater variability due to the ambiguity of the task. Under this setting, the relation between annotators' variability and model performance has received little attention. We present a large-scale study on the variability of cell annotations among 120 board-certified pathologists and how it affects the performance of a deep learning model. We propose a method to measure such variability, and by excluding those annotators with low variability, we verify the trade-off between the amount of data and its quality. We found that naively increasing the data size at the expense of inter-rater variability does not necessarily lead to better-performing models in cell detection. Instead, decreasing the inter-rater variability with the expense of decreasing dataset size increased the model performance. Furthermore, models trained from data annotated with lower inter-labeler variability outperform those from higher inter-labeler variability. These findings suggest that the evaluation of the annotators may help tackle the fundamental budget issues in the histopathology domain
Code (0)
등록된 구현이 없습니다.
Tasks
Cell DetectionSimilar Papers 제목 키워드 기반
Reliability of deep learning models for anatomical landmark detection: The role of inter-rater variability
Automated detection of anatomical landmarks plays a crucial role in many diagnostic and surgical applications. Progresses in deep learning (DL) methods have resulted in significant performance enhancement in tasks relate…
Anatomical Landmark DetectionDiagnosticHow inter-rater variability relates to aleatoric and epistemic uncertainty: a case study with deep learning-based paraspinal muscle segmentation
Recent developments in deep learning (DL) techniques have led to great performance improvement in medical image segmentation tasks, especially with the latest Transformer model and its variants. While labels from fusing …
Image SegmentationMedical Image SegmentationSemantic SegmentationLabel fusion and training methods for reliable representation of inter-rater uncertainty
Medical tasks are prone to inter-rater variability due to multiple factors such as image quality, professional experience and training, or guideline clarity. Training deep learning networks with annotations from multiple…
SegmentationQUBIQ: Uncertainty Quantification for Biomedical Image Segmentation Challenge
Uncertainty in medical image segmentation tasks, especially inter-rater variability, arising from differences in interpretations and annotations by various experts, presents a significant challenge in achieving consisten…
Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation+1Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation
Few-shot medical image segmentation methods typically assume a single ground-truth annotation, overlooking systematic variability across expert raters commonly observed in clinical datasets. We propose an attention-based…
Medical Image SegmentationPersonalized Segmentation