paper-with-me

홈 › Papers

Domain-Adaptive Pre-training of Self-Supervised Foundation Models for Medical Image Classification in Gastrointestinal Endoscopy

2024-10-21 · Marcel Roth, Micha V. Nowak, Adrian Krenzer, Frank Puppe

Video capsule endoscopy has transformed gastrointestinal endoscopy (GIE) diagnostics by offering a non-invasive method for capturing detailed images of the gastrointestinal tract, enabling early disease detection. However, its potential is limited by the sheer volume of images generated during the imaging procedure, which can take anywhere from 6-8 hours and often produce up to 1 million images, necessitating automated analysis. Additionally, the variability of these images, combined with the need for expert annotations and the scarcity of large, high-quality labeled datasets, constrains the effectiveness of current medical image analysis models. To address this, we introduce a novel large GIE dataset, called EndoExtend24, created by merging ten existing public and private datasets, ensuring patient integrity across splits. EndoExtend24 includes over 226,000 labeled images, as well as dynamic class mappings, which allow unified training across datasets with differing labeling granularity, supporting up to 123 distinct pathological findings. Further, we propose to leverage domain adaptive pre-training of foundation models trained with self-supervision on generic image data, to adapt them to the task of GIE medical image diagnosis. Specifically, the EVA-02 model, which is based on the ViT architecture and trained on ImageNet-22k with masked image modeling (using EVA-CLIP as a MIM teacher), is pre-trained on the EndoExtend24 dataset to achieve domain adaptation, and finally trained on the Capsule Endoscopy 2024 Challenge dataset. Our model demonstrates robust performance, securing third place in the Capsule Endoscopy 2024 Challenge. We achieved a macro AUC of 0.762 and a balanced accuracy of 37.1% on the test set. These results emphasize the effectiveness of our domain-adaptive pre-training approach and the enriched EndoExtend24 dataset in advancing gastrointestinal endoscopy diagnostics.

📄 PDF Abstract BibTeX arXiv:2410.21302

Code (1)

mvrcii/capsule_vision_challenge_2024 pytorch

Tasks

Domain Adaptationimage-classificationImage ClassificationMedical DiagnosisMedical Image AnalysisMedical Image Classification

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
MIM 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Unsupervised Domain Adaptive Re-Identification: Theory and Practice

2018-07-30 · Liangchen Song, Cheng Wang, Lefei Zhang, Bo Du 외

We study the problem of unsupervised domain adaptive re-identification (re-ID) which is an active topic in computer vision but lacks a theoretical foundation. We first extend existing unsupervised domain adaptive classif…

General ClassificationUnsupervised Domain Adaptation

Efficient Domain-Adaptive Multi-Task Dense Prediction with Vision Foundation Models

2025-09-28 · Beomseok Kang, Niluthpol Chowdhury Mithun, Mikhail Sizintsev, Han-Pang Chiu 외 arxiv

Multi-task dense prediction, which aims to jointly solve tasks like semantic segmentation and depth estimation, is crucial for robotics applications but suffers from domain shift when deploying models in new environments…

Unsupervised Domain AdaptationSemantic SegmentationMulti-Task LearningDepth Estimation

DA-SSL: self-supervised domain adaptor to leverage foundational models in turbt histopathology slides

2025-12-15 · Haoyue Zhang, Meera Chappidi, Erolcan Sayar, Helen Richards 외 arxiv

Recent deep learning frameworks in histopathology, particularly multiple instance learning (MIL) combined with pathology foundational models (PFMs), have shown strong performance. However, PFMs exhibit limitations on cer…

Multiple Instance LearningDomain Adaptation

Person Re-ID in 2025: Supervised, Self-Supervised, and Language-Aligned. What Works?

2026-01-28 · Lakshman Balasubramanian arxiv

Person Re-Identification (ReID) remains a challenging problem in computer vision. This work reviews various training paradigm and evaluates the robustness of state-of-the-art ReID models in cross-domain applications and …

Person Re-Identification

CLEF: Clinically-Guided Contrastive Learning for Electrocardiogram Foundation Models

2025-12-01 · Yuxuan Shu, Peter H. Charlton, Fahim Kawsar, Jussi Hernesniemi 외 arxiv

The electrocardiogram (ECG) is a key diagnostic tool in cardiovascular health. Single-lead ECG recording is integrated into both clinical-grade and consumer wearables. While self-supervised pretraining of foundation mode…

Self-Supervised LearningContrastive Learning