paper-with-me

Papers

T3D: Advancing 3D Medical Vision-Language Pre-training by Learning Multi-View Visual Consistency

2023-12-03 · Che Liu, Cheng Ouyang, Yinda Chen, Cesar César Quilodrán-Casas, Lei Ma, Jie Fu, Yike Guo, Anand Shah, Wenjia Bai, Rossella Arcucci

While 3D visual self-supervised learning (vSSL) shows promising results in capturing visual representations, it overlooks the clinical knowledge from radiology reports. Meanwhile, 3D medical vision-language pre-training (MedVLP) remains underexplored due to the lack of a large-scale, publicly available 3D medical image-report dataset. To bridge this gap, we introduce CT-3DVLP, the first and largest public 3D volume-report dataset, establishing a comprehensive benchmark for 3D MedVLP research. Meanwhile, we propose the T3D framework, which enhances 3D MedVLP beyond naive CLIP-style alignment that directly pairs volumes with reports but neglects local visual representations. Instead, we introduce Text-informed Multi-view Alignment (TMA), a novel approach that clusters volumetric data while enforcing consistency across different views of the same volume-report pair. TMA integrates textual features into fine-grained visual representations, ensuring contextual coherence across views. We evaluate T3D across multiple downstream tasks in both unimodal and cross-modal settings, including zero-shot and fine-tuned classification, cross-modal retrieval, report generation, and semantic segmentation. Our results show that T3D consistently outperforms existing vSSL and multimodal methods, demonstrating superior zero-shot and fine-tuning capabilities and setting a new benchmark for 3D medical image understanding.

📄 PDF Abstract BibTeX arXiv:2312.01529

Code (0)

등록된 구현이 없습니다.

Tasks

Clinical KnowledgeContrastive LearningCross-Modal RetrievalImage RestorationMedical Image AnalysisRepresentation LearningSelf-Supervised LearningSemantic SegmentationTumor Segmentation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Improving Medical Visual Representations via Radiology Report Generation

2023-10-30 · Keegan Quigley, Miriam Cha, Josh Barua, Geeticka Chauhan 외

Vision-language pretraining has been shown to produce high-quality visual encoders which transfer efficiently to downstream computer vision tasks. Contrastive learning approaches have increasingly been adopted for medica…

Contrastive LearningDecoderImage CaptioningMedical Image Analysis

Advancing Medical Representation Learning Through High-Quality Data

2025-03-18 · Negin Baghbanzadeh, Adibvafa Fallahpour, Yasaman Parhizkar, Franklin Ogidi 외

Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medical dataset from PubMed Central, contain…

Representation Learningzero-shot-classificationZero-Shot Learning

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

2025-09-29 · Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen, Trung Thanh Nguyen 외 arxiv

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in gene…

Data Augmentation

Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making

2025-08-08 · Kaitao Chen, Mianxin Liu, Daoming Zong, Chaoyue Ding 외 arxiv

Complex medical decision-making involves cooperative workflows operated by different clinicians. Designing AI multi-agent systems can expedite and augment human-level clinical decision-making. Existing multi-agent resear…

Instruction FollowingQuestion Answering

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

2025-01-13 · Difei Gu, Yunhe Gao, Yang Zhou, Mu Zhou 외

Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinical workflow. Current approaches either fo…

Concept AlignmentImage CaptioningRetrieval-augmented Generation