Label-Assemble: Leveraging Multiple Datasets with Partial Labels
The success of deep learning relies heavily on large labeled datasets, but we often only have access to several small datasets associated with partial labels. To address this problem, we propose a new initiative, "Label-Assemble", that aims to unleash the full potential of partial labels from an assembly of public datasets. We discovered that learning from negative examples facilitates both computer-aided disease diagnosis and detection. This discovery will be particularly crucial in novel disease diagnosis, where positive examples are hard to collect, yet negative examples are relatively easier to assemble. For example, assembling existing labels from NIH ChestX-ray14 (available since 2017) significantly improves the accuracy of COVID-19 diagnosis from 96.3% to 99.3%. In addition to diagnosis, assembling labels can also improve disease detection, e.g., the detection of pancreatic ductal adenocarcinoma (PDAC) can greatly benefit from leveraging the labels of Cysts and PanNets (two other types of pancreatic abnormalities), increasing sensitivity from 52.1% to 84.0% while maintaining a high specificity of 98.0%.
Code (2)
Tasks
COVID-19 DiagnosisSpecificityMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
PSScreen: Partially Supervised Multiple Retinal Disease Screening
Leveraging multiple partially labeled datasets to train a model for multiple retinal disease screening reduces the reliance on fully annotated datasets, but remains challenging due to significant domain shifts across tra…
Domain GeneralizationHeterogeneous Risk Minimization
Machine learning algorithms with empirical risk minimization usually suffer from poor generalization performance due to the greedy exploitation of correlations among the training data, which are not stable under distribu…
Leveraging knowledge distillation for partial multi-task learning from multiple remote sensing datasets
Partial multi-task learning where training examples are annotated for one of the target tasks is a promising idea in remote sensing as it allows combining datasets annotated for different tasks and predicting more tasks …
Knowledge DistillationMulti-Task Learningobject-detectionObject Detection+1Marginal loss and exclusion loss for partially supervised multi-organ segmentation
Annotating multiple organs in medical images is both costly and time-consuming; therefore, existing multi-organ datasets with labels are often low in sample size and mostly partially labeled, that is, a dataset has a few…
Organ SegmentationSegmentationSegregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
This paper proposes a segregated temporal assembly recurrent (STAR) network for weakly-supervised multiple action detection. The model learns from untrimmed videos with only supervision of video-level labels and makes pr…
Action DetectionMultiple Action Detection