paper-with-me

홈 › Papers

Not again! Data Leakage in Digital Pathology

2019-09-14 · Nicole Bussola, Alessia Marcolini, Valerio Maggio, Giuseppe Jurman, Cesare Furlanello

Bioinformatics of high throughput omics data (e.g. microarrays and proteomics) has been plagued by uncountable issues with reproducibility at the start of the century. Concerns have motivated international initiatives such as the FDA's led MAQC Consortium, addressing reproducibility of predictive biomarkers by means of appropriate Data Analysis Plans (DAPs). For instance, repreated cross-validation is a standard procedure meant at mitigating the risk that information from held-out validation data may be used during model selection. We prove here that, many years later, Data Leakage can still be a non-negligible overfitting source in deep learning models for digital pathology. In particular, we evaluate the impact of (i) the presence of multiple images for each subject in histology collections; (ii) the systematic adoption of training over collection of subregions (i.e. "tiles" or "patches") extracted for the same subject. We verify that accuracy scores may be inflated up to 41%, even if a well-designed 10x5 iterated cross-validation DAP is applied, unless all images from the same subject are kept together either in the internal training or validation splits. Results are replicated for 4 classification tasks in digital pathology on 3 datasets, for a total of 373 subjects, and 543 total slides (around 27, 000 tiles). Impact of applying transfer learning strategies with models pre-trained on general-purpose or digital pathology datasets is also discussed.

📄 PDF Abstract BibTeX arXiv:1909.06539

Code (0)

등록된 구현이 없습니다.

Tasks

Model SelectionTransfer Learning

Similar Papers 제목 키워드 기반

RCdpia: A Renal Carcinoma Digital Pathology Image Annotation dataset based on pathologists

2024-03-17 · Qingrong Sun, Weixiang Zhong, Jie zhou, Chong Lai 외

The annotation of digital pathological slide data for renal cell carcinoma is of paramount importance for correct diagnosis of artificial intelligence models due to the heterogeneous nature of the tumor. This process not…

Improved statistical benchmarking of digital pathology models using pairwise frames evaluation

2023-06-07 · Ylaine Gerardin, John Shamshoian, Judy Shen, Nhat Le 외

Nested pairwise frames is a method for relative benchmarking of cell or tissue digital pathology models against manual pathologist annotations on a set of sampled patches. At a high level, the method compares agreement b…

BenchmarkingClassification

DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset

2026-05-05 · Carlijn Lems, Sander Moonemans, Natálie Klubíčková, Biagio Brattoli 외 arxiv

Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent benchmarking to assess its potential in assisting pathologists in routi…

Visual Question AnsweringAnswer Generation

Learned Image Compression and Restoration for Digital Pathology

2025-03-31 · SeonYeong Lee, EonSeung Seong, Dongeon Lee, SiYeoul Lee 외

Digital pathology images play a crucial role in medical diagnostics, but their ultra-high resolution and large file sizes pose significant challenges for storage, transmission, and real-time visualization. To address the…

DiagnosticImage CompressionImage Reconstructionwhole slide images

Foundation Models and Information Retrieval in Digital Pathology

2024-03-13 · H. R. Tizhoosh

The paper reviews the state-of-the-art of foundation models, LLMs, generative AI, information retrieval and CBIR in digital pathology

Information RetrievalRetrieval