paper-with-me

홈 › Papers

Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation

2024-12-10 · Qianhao Han, Junyi Liu, Zengchang Qin, Zheng Zheng

Automating radiology report generation can significantly reduce the workload of radiologists and enhance the accuracy, consistency, and efficiency of clinical documentation.We propose a novel cross-modal framework that uses MedCLIP as both a vision extractor and a retrieval mechanism to improve the process of medical report generation.By extracting retrieved report features and image features through an attention-based extract module, and integrating them with a fusion module, our method improves the coherence and clinical relevance of generated reports.Experimental results on the widely used IU-Xray dataset demonstrate the effectiveness of our approach, showing improvements over commonly used methods in both report quality and relevance.Additionally, ablation studies provide further validation of the framework, highlighting the importance of accurate report retrieval and feature integration in generating comprehensive medical reports.

📄 PDF Abstract BibTeX arXiv:2412.07141

Code (1)

QianhaoHan/Cross-Modality-Medical-Report-Generation 공식 구현 pytorch

Tasks

Retrieval

Similar Papers 제목 키워드 기반

A Multimodal Approach For Endoscopic VCE Image Classification Using BiomedCLIP-PubMedBERT

2024-10-25 · Nagarajan Ganapathy, Podakanti Satyajith Chary, Teja Venkata Ramana Kumar Pithani, Pavan Kavati 외

This Paper presents an advanced approach for fine-tuning BiomedCLIP PubMedBERT, a multimodal model, to classify abnormalities in Video Capsule Endoscopy (VCE) frames, aiming to enhance diagnostic efficiency in gastrointe…

Diagnosticimage-classificationImage ClassificationLanguage Modeling+1

Multi-task Cross-modal Learning for Chest X-ray Image Retrieval

2026-01-08 · Zhaohui Liang, Sivaramakrishnan Rajaraman, Niccolo Marini, Zhiyun Xue 외 arxiv

CLIP and BiomedCLIP are examples of vision-language foundation models and offer strong cross-modal embeddings; however, they are not optimized for fine-grained medical retrieval tasks, such as retrieving clinically relev…

Cross-Modal RetrievalMulti-Task LearningImage RetrievalText Retrieval

MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentation

2026-02-23 · Taha Koleilat, Hojat Asgariandehkordi, Omid Nejati Manzari, Berardino Barile 외 arxiv

Medical image segmentation remains challenging due to limited annotations for training, ambiguous anatomical features, and domain shifts. While vision-language models such as CLIP offer strong cross-modal representations…

Medical Image Segmentation

Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

2021-12-27 · Sedigheh Eslami, Gerard de Melo, Christoph Meinel

Contrastive Language--Image Pre-training (CLIP) has shown remarkable success in learning with cross-modal supervision from extensive amounts of image--text pairs collected online. Thus far, the effectiveness of CLIP has …

ArticlesMedical Visual Question AnsweringMeta-LearningQuestion Answering+3

BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

2023-03-02 · Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu 외

Biomedical data is inherently multimodal, comprising physical measurements and natural language narratives. A generalist biomedical AI model needs to simultaneously process different modalities of data, including text an…

ArticlesMedical Visual Question AnsweringPneumonia DetectionQuestion Answering+3