paper-with-me

홈 › Papers

DALE-CT: Depth-Aware Foundation Models for Computed Tomography

2026-06-05 · Evan W. Damron, Mahmut S. Gokmen, Mitchell A. Klusty, Caroline N. Leach, Emily B. Collier, V. K. Cody Bumgardner arxiv

Recent breakthroughs in self-supervised learning (SSL), such as the Latent-Euclidean Joint-Embedding Predictive Architecture (LeJEPA), alongside successes in integrating visual encoders with language models, have driven the demand for adaptable, high-capacity vision encoders in Computed Tomography (CT). In this work, we explore 2D slice-based architectures as a flexible alternative to native 3D models for processing volumetric CT data. Using the CT-RATE dataset, we trained DALE-CT (Depth-Aware Latent-Euclidean Computed Tomography), a 2D model family built entirely from scratch using LeJEPA, and compared its performance against a continually pre-trained DINOv2 baseline. To enhance representation quality, we developed a novel 3D depth-aware pre-training strategy anchored by dense auxiliary supervision from both automated anatomical masks and human-annotated abnormalities. Under linear probe evaluation with Multiple Instance Learning (MIL) for multi-abnormality detection, the frozen backbone of this dual-supervised model (DALE-CT-2S) achieves a Macro AUROC of 0.833. This performance demonstrates near-parity with state-of-the-art 3D vision-language models, achieved entirely from scratch with significantly less data and no textual supervision. To ensure reproducibility, all training code, evaluation scripts, and model weights have been made publicly available.

📄 PDF Abstract BibTeX arXiv:2606.07775

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple Instance LearningSelf-Supervised Learning

Similar Papers 제목 키워드 기반

TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models

2025-11-30 · Tim Veenboer, George Yiasemis, Eric Marcus, Vivien Van Veldhuizen 외 arxiv

Existing foundation models (FMs) in the medical domain often require extensive fine-tuning or rely on training resource-intensive decoders, while many existing encoders are pretrained with objectives biased toward specif…

Probabilistic Dalek -- Emulator framework with probabilistic prediction for supernova tomography

2022-09-20 · Wolfgang Kerzendorf, Nutan Chen, Jack O'Brien, Johannes Buchner 외

Supernova spectral time series can be used to reconstruct a spatially resolved explosion model known as supernova tomography. In addition to an observed spectral time series, a supernova tomography requires a radiative t…

Active LearningCPUTime SeriesTime Series Analysis+1

Vision Foundation Models for Computed Tomography

2025-01-15 · Suraj Pai, Ibrahim Hadzic, Dennis Bontempi, Keno Bressem 외

Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed e…

Computed Tomography (CT)Contrastive LearningImage RetrievalMedical Image Retrieval+1

Toward a Foundation Plug-and-Play Prior for Computed Tomography Reconstruction via a Multimodal Diffusion Model

2026-08-24 · Haley Duba-Sullivan, Patxi Fernandez-Zelaia, Obaidullah Rahman, Amirkoushyar Ziabari arxiv

Computed tomography (CT) throughput is limited by scan time, which grows with both the number of projections acquired and the detector integration time for each. Reconstructing high-quality volumes from sparse-view or lo…

Ctcovid19: Automatic Covid-19 Model for Computed Tomography Scans Using Deep Learning

2024-08-09 · Based Intelligence Medicine 2024 8 · Carlos Antunes, João Rodrigues, António Cunha

Summary COVID-19 is an extremely contagious respiratory sickness instigated by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Common symptoms encompass fever, cough, fatigue, and breathing difficulties…