Dataset creation for supervised deep learning-based analysis of microscopic images -- review of important considerations and recommendations
Supervised deep learning (DL) receives great interest for automated analysis of microscopic images with an increasing body of literature supporting its potential. The development and validation of those DL models relies heavily on the availability of high-quality, large-scale datasets. However, creating such datasets is a complex and resource-intensive process, often hindered by challenges such as time constraints, domain variability, and risks of bias in image collection and label creation. This review provides a comprehensive guide to the critical steps in dataset creation, including: 1) image acquisition, 2) selection of annotation software, and 3) annotation creation. In addition to ensuring a sufficiently large number of images, it is crucial to address sources of image variability (domain shifts) - such as those related to slide preparation and digitization - that could lead to algorithmic errors if not adequately represented in the training data. Key quality criteria for annotations are the three "C"s: correctness, completeness, and consistency. This review explores methods to enhance annotation quality through the use of advanced techniques that mitigate the limitations of single annotators. To support dataset creators, a standard operating procedure (SOP) is provided as supplemental material, outlining best practices for dataset development. Furthermore, the article underscores the importance of open datasets in driving innovation and enhancing reproducibility of DL research. By addressing the challenges and offering practical recommendations, this review aims to advance the creation of and availability to high-quality, large-scale datasets, ultimately contributing to the development of generalizable and robust DL models for pathology applications.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Optimizations of Autoencoders for Analysis and Classification of Microscopic In Situ Hybridization Images
Currently, analysis of microscopic In Situ Hybridization images is done manually by experts. Precise evaluation and classification of such microscopic images can ease experts' work and reveal further insights about the d…
Deep LearningSemantic Aware Data Augmentation for Cell Nuclei Microscopical Images With Artificial Neural Networks
There exists many powerful architectures for object detection and semantic segmentation of both biomedical and natural images. However, a difficulty arises in the ability to create training datasets that are large an…
Data Augmentationobject-detectionObject DetectionSegmentation+1Unsupervised Representations of Pollen in Bright-Field Microscopy
We present the first unsupervised deep learning method for pollen analysis using bright-field microscopy. Using a modest dataset of 650 images of pollen grains collected from honey, we achieve family level identification…
ClusteringClassification Beats Regression: Counting of Cells from Greyscale Microscopic Images based on Annotation-free Training Samples
Modern methods often formulate the counting of cells from microscopic images as a regression problem and more or less rely on expensive, manually annotated training images (e.g., dot annotations indicating the centroids …
Data Augmentationimage-classificationImage ClassificationAutomatic microscopic cell counting by use of unsupervised adversarial domain adaptation and supervised density regression
Accurate cell counting in microscopic images is important for medical diagnoses and biological studies. However, manual cell counting is very time-consuming, tedious, and prone to subjective errors. We propose a new dens…
Automatic Cell CountingDomain Adaptationregression