paper-with-me

Papers

Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis

2025-01-07 · Haoran Lai, Zihang Jiang, Qingsong Yao, Rongsheng Wang, ZhiYang He, Xiaodong Tao, Wei Wei, Weifu Lv, S. Kevin Zhou

3D medical images such as Computed tomography (CT) are widely used in clinical practice, offering a great potential for automatic diagnosis. Supervised learning-based approaches have achieved significant progress but rely heavily on extensive manual annotations, limited by the availability of training data and the diversity of abnormality types. Vision-language alignment (VLA) offers a promising alternative by enabling zero-shot learning without additional annotations. However, we empirically discover that the visual and textural embeddings after alignment endeavors from existing VLA methods form two well-separated clusters, presenting a wide gap to be bridged. To bridge this gap, we propose a Bridged Semantic Alignment (BrgSA) framework. First, we utilize a large language model to perform semantic summarization of reports, extracting high-level semantic information. Second, we design a Cross-Modal Knowledge Interaction (CMKI) module that leverages a cross-modal knowledge bank as a semantic bridge, facilitating interaction between the two modalities, narrowing the gap, and improving their alignment. To comprehensively evaluate our method, we construct a benchmark dataset that includes 15 underrepresented abnormalities as well as utilize two existing benchmark datasets. Experimental results demonstrate that BrgSA achieves state-of-the-art performances on both public benchmark datasets and our custom-labeled dataset, with significant improvements in zero-shot diagnosis of underrepresented abnormalities.

📄 PDF Abstract BibTeX arXiv:2501.03565

Code (0)

등록된 구현이 없습니다.

Tasks

Computed Tomography (CT)Large Language ModelZero-Shot Learning

Similar Papers 제목 키워드 기반

TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models

2025-06-13 · Ziyang Luo, Nian Liu, Xuguang Yang, Salman Khan 외

Audio-Visual Segmentation (AVS) faces a fundamental challenge of effectively aligning audio and visual modalities. While recent approaches leverage foundation models to address data scarcity, they often rely on single-mo…

cross-modal alignmentSegmentation

CARZero: Cross-Attention Alignment for Radiology Zero-Shot Classification

2024-02-27 · CVPR 2024 1 · Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang 외

The advancement of Zero-Shot Learning in the medical domain has been driven forward by using pre-trained models on large-scale image-text pairs, focusing on image-text alignment. However, existing methods primarily rely …

ClassificationDiagnosticLanguage ModellingLarge Language Model+3

LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?

2025-03-10 · Bangyan Li, Wenxuan Huang, Yunhang Shen, Yeqiang Wang 외

Recently, multimodal large models (MLLMs) have demonstrated exceptional capabilities in visual understanding and reasoning across various vision-language tasks. However, MLLMs usually perform poorly in zero-shot medical …

cross-modal alignment

BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

2025-08-06 · Tongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui Liu arxiv

Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with amb…

Zero-shot GeneralizationStereo Depth Estimation

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability

2025-04-10 · Jonggwon Park, Soobum Kim, Byungmu Yoon, Kyoyun Choi

Recent advancements in multi-modal models have significantly improved vision-language alignment in radiology. However, existing approaches struggle to effectively utilize complex radiology reports for learning, rely on l…

Contrastive LearningOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentation+3