Reverse Classification Accuracy: Predicting Segmentation Performance in the Absence of Ground Truth
When integrating computational tools such as automatic segmentation into clinical practice, it is of utmost importance to be able to assess the level of accuracy on new data, and in particular, to detect when an automatic method fails. However, this is difficult to achieve due to absence of ground truth. Segmentation accuracy on clinical data might be different from what is found through cross-validation because validation data is often used during incremental method development, which can lead to overfitting and unrealistic performance expectations. Before deployment, performance is quantified using different metrics, for which the predicted segmentation is compared to a reference segmentation, often obtained manually by an expert. But little is known about the real performance after deployment when a reference is unavailable. In this paper, we introduce the concept of reverse classification accuracy (RCA) as a framework for predicting the performance of a segmentation method on new data. In RCA we take the predicted segmentation from a new image to train a reverse classifier which is evaluated on a set of reference images with available ground truth. The hypothesis is that if the predicted segmentation is of good quality, then the reverse classifier will perform well on at least some of the reference images. We validate our approach on multi-organ segmentation with different classifiers and segmentation methods. Our results indicate that it is indeed possible to predict the quality of individual segmentations, in the absence of ground truth. Thus, RCA is ideal for integration into automatic processing pipelines in clinical routine and as part of large-scale image analysis studies.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationOrgan SegmentationSegmentationSimilar Papers 제목 키워드 기반
Automated Quality Control in Image Segmentation: Application to the UK Biobank Cardiac MR Imaging Study
Background: The trend towards large-scale studies including population imaging poses new challenges in terms of quality control (QC). This is a particular issue when automatic processing tools, e.g. image segmentation me…
General ClassificationImage SegmentationSegmentationSemantic SegmentationDomain Adaptation for MRI Organ Segmentation using Reverse Classification Accuracy
The variations in multi-center data in medical imaging studies have brought the necessity of domain adaptation. Despite the advancement of machine learning in automatic segmentation, performance often degrades when algor…
ClassificationDomain AdaptationGeneral ClassificationOrgan Segmentation+2To Annotate or Not? Predicting Performance Drop under Domain Shift
Performance drop due to domain-shift is an endemic problem for NLP models in production. This problem creates an urge to continuously annotate evaluation datasets to measure the expected drop in the model performance whi…
General ClassificationPOSPOS TaggingSentiment Analysis+1Reverse Multi-Label Learning
Multi-label classification is the task of predicting potentially multiple labels for a given instance. This is common in several applications such as image annotation, document classification and gene function prediction…
ClassificationDocument ClassificationGeneral ClassificationMulti-Label Classification+3Dense Pixel-Labeling for Reverse-Transfer and Diagnostic Learning on Lung Ultrasound for COVID-19 and Pneumonia Detection
We propose using a pre-trained segmentation model to perform diagnostic classification in order to achieve better generalization and interpretability, terming the technique reverse-transfer learning. We present an archit…
ClassificationDiagnosticPneumonia DetectionSegmentation+1