paper-with-me

홈 › Papers

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

2025-08-16 · Ziyang Zhang, Yang Yu, Xulei Yang, Si Yong Yeo arxiv

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating large-scale paired data in the medical field for volumetric modalities such as CT scans remains a challenging and time-intensive process. This difficulty often limits the performance on downstream tasks. To address these challenges, we propose a novel vision-language pre-training (VLP) framework, termed as \textbf{VELVET-Med}, specifically designed for limited volumetric data such as 3D CT and associated radiology reports. Instead of relying on large-scale data collection, our method focuses on the development of effective pre-training objectives and model architectures. The key contributions are: 1) We incorporate uni-modal self-supervised learning into VLP framework, which are often underexplored in the existing literature. 2) We propose a novel language encoder, termed as \textbf{TriBERT}, for learning multi-level textual semantics. 3) We devise the hierarchical contrastive learning to capture multi-level vision-language correspondence. Using only 38,875 scan-report pairs, our approach seeks to uncover rich spatial and semantic relationships embedded in volumetric medical images and corresponding clinical narratives, thereby enhancing the generalization ability of the learned encoders. The resulting encoders exhibit strong transferability, achieving state-of-the-art performance across a wide range of downstream tasks, including 3D segmentation, cross-modal retrieval, visual question answering, and report generation.

📄 PDF Abstract BibTeX arXiv:2508.12108

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringSelf-Supervised LearningCross-Modal RetrievalContrastive Learning

Similar Papers 제목 키워드 기반

Velvet Worms (Onychophora) in Folklore and Art: Geographic Pattern, Types of Cultural Reference and Public Perception

2016-05-21

Aims: To document and preserve folkloric beliefs and art inspired by velvet worms (Onychophora), rare invertebrates that are considered "living fossils", have full placental organs and capture prey with a rough "net" bui…

Cultural Vocal Bursts Intensity PredictionImage Retrieval

VELVET: a noVel Ensemble Learning approach to automatically locate VulnErable sTatements

2021-12-20 · Yangruibo Ding, Sahil Suneja, Yunhui Zheng, Jim Laredo 외

Automatically locating vulnerable statements in source code is crucial to assure software security and alleviate developers' debugging efforts. This becomes even more important in today's software ecosystem, where vulner…

Ensemble Learning

Non-Exponential Reverberation Modeling Using Dark Velvet Noise

2024-03-29 · Jon Fagerström, Sebastian J. Schelcht, Vesa Välimäki

Previous research on late-reverberation modeling has mainly focused on exponentially decaying room impulse responses, whereas methods for accurately modeling non-exponential reverberation remain challenging. This paper e…

Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers

2025-11-21 · Cris Claessens, Christiaan Viviers, Giacomo D'Amicantonio, Egor Bondarev 외 arxiv

We introduce SPECTRE, a fully transformer-based foundation model for volumetric computed tomography (CT). Our Self-Supervised & Cross-Modal Pretraining for CT Representation Extraction (SPECTRE) approach utilizes scalabl…

Contrastive Learning

Training-Free Zero-Shot Anomaly Detection in 3D Brain MRI with 2D Foundation Models

2026-02-17 · Tai Le-Gia, Jaehyun Ahn arxiv

Zero-shot anomaly detection (ZSAD) has gained increasing attention in medical imaging as a way to identify abnormalities without task-specific supervision, but most advances remain limited to 2D datasets. Extending ZSAD …

Anomaly Detection