paper-with-me

Papers

Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology

2025-03-24 · Boqi Chen, Cédric Vincent-Cuaz, Lydia A. Schoenpflug, Manuel Madeira, Lisa Fournier, Vaishnavi Subramanian, Sonali Andani, Samuel Ruiperez-Campillo, Julia E. Vogt, Raphaëlle Luisier, Dorina Thanou, Viktor H. Koelzer, Pascal Frossard, Gabriele Campanella, Gunnar Rätsch

Vision foundation models (FMs) are accelerating the development of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size, diversity, and balance of the pre-training data. However, data selection has been primarily guided by expert knowledge at the WSI level, focusing on factors such as disease classification and tissue types, while largely overlooking the granular details available at the tile level. In this paper, we investigate the potential of unsupervised automatic data curation at the tile-level, taking into account 350 million tiles. Specifically, we apply hierarchical clustering trees to pre-extracted tile embeddings, allowing us to sample balanced datasets uniformly across the embedding space of the pretrained FM. We further identify these datasets are subject to a trade-off between size and balance, potentially compromising the quality of representations learned by FMs, and propose tailored batch sampling strategies to mitigate this effect. We demonstrate the effectiveness of our method through improved performance on a diverse range of clinically relevant downstream tasks.

📄 PDF Abstract BibTeX arXiv:2503.18709

Code (0)

등록된 구현이 없습니다.

Tasks

whole slide images

Similar Papers 제목 키워드 기반

FoundationStereo: Zero-Shot Stereo Matching

2025-01-17 · CVPR 2025 1 · Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz 외

Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization - a hallmark of foundation models in other compu…

Depth EstimationDiversityStereo Depth EstimationStereo Matching+1

OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation

2025-09-22 · Paweł Budzianowski, Emilia Wiśnios, Michał Tyrolski, Gracjan Góral 외 arxiv

Data scarcity remains one of the most limiting factors in driving progress in robotics. However, the amount of available robotics data in the wild is growing exponentially, creating new opportunities for large-scale data…

Efficient Curation of Invertebrate Image Datasets Using Feature Embeddings and Automatic Size Comparison

2024-12-20 · Mikko Impiö, Philipp M. Rehsen, Jenni Raitoharju

The amount of image datasets collected for environmental monitoring purposes has increased in the past years as computer vision assisted methods have gained interest. Computer vision applications rely on high-quality dat…

Outlier Detection

Large Language Models as Automated Aligners for benchmarking Vision-Language Models

2023-11-24 · Yuanfeng Ji, Chongjian Ge, Weikai Kong, Enze Xie 외

With the advancements in Large Language Models (LLMs), Vision-Language Models (VLMs) have reached a new level of sophistication, showing notable competence in executing intricate cognition and reasoning tasks. However, e…

BenchmarkingWorld Knowledge

External Evaluation of Event Extraction Classifiers for Automatic Pathway Curation: An extended study of the mTOR pathway

2017-07-07 · WS 2017 8 · Wojciech Kusa, Michael Spranger

This paper evaluates the impact of various event extraction systems on automatic pathway curation using the popular mTOR pathway. We quantify the impact of training data sets as well as different machine learning classif…

BIG-bench Machine LearningEvent Extraction