paper-with-me

홈 › Papers

Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation

2025-05-02 · Daniele Molino, Francesco Di Feola, Linlin Shen, Paolo Soda, Valerio Guarrasi

Generative models have revolutionized Artificial Intelligence (AI), particularly in multimodal applications. However, adapting these models to the medical domain poses unique challenges due to the complexity of medical data and the stringent need for clinical accuracy. In this work, we introduce a framework specifically designed for multimodal medical data generation. By enabling the generation of multi-view chest X-rays and their associated clinical report, it bridges the gap between general-purpose vision-language models and the specialized requirements of healthcare. Leveraging the MIMIC-CXR dataset, the proposed framework shows superior performance in generating high-fidelity images and semantically coherent reports. Our quantitative evaluation reveals significant results in terms of FID and BLEU scores, showcasing the quality of the generated data. Notably, our framework achieves comparable or even superior performance compared to real data on downstream disease classification tasks, underlining its potential as a tool for medical research and diagnostics. This study highlights the importance of domain-specific adaptations in enhancing the relevance and utility of generative models for clinical applications, paving the way for future advancements in synthetic multimodal medical data generation.

📄 PDF Abstract BibTeX arXiv:2505.01091

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

MicarVLMoE: A Modern Gated Cross-Aligned Vision-Language Mixture of Experts Model for Medical Image Captioning and Report Generation

2025-04-29 · Amaan Izhar, Nurul Japar, Norisma Idris, Ting Dang

Medical image reporting (MIR) aims to generate structured clinical descriptions from radiological images. Existing methods struggle with fine-grained feature extraction, multimodal alignment, and generalization across di…

cross-modal alignmentDecoderImage CaptioningMixture-of-Experts

Neuradicon: operational representation learning of neuroimaging reports

2021-07-21 · Henry Watkins, Robert Gray, Adam Julius, Yee-Haur Mah 외

Radiological reports typically summarize the content and interpretation of imaging studies in unstructured form that precludes quantitative analysis. This limits the monitoring of radiological services to throughput undi…

FormLanguage ModellingRepresentation Learning

Vision-Language Modelling For Radiological Imaging and Reports In The Low Data Regime

2023-03-30 · Rhydian Windsor, Amir Jamaludin, Timor Kadir, Andrew Zisserman

This paper explores training medical vision-language models (VLMs) -- where the visual and language inputs are embedded into a common space -- with a particular focus on scenarios where training data is limited, as is of…

Image RetrievalLanguage ModellingRetrieval

3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models

2024-09-28 · Hao Chen, Wei Zhao, Yingli Li, Tianyang Zhong 외

Medical image analysis is crucial in modern radiological diagnostics, especially given the exponential growth in medical imaging data. The demand for automated report generation systems has become increasingly urgent. Wh…

DiagnosticLanguage ModelingLanguage ModellingMedical Image Analysis+3

SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models

2024-04-27 · Manav Nitin Kapadnis, Sohan Patnaik, Abhilash Nandy, Sourjyadip Ray 외

Radiology Report Generation (R2Gen) demonstrates how Multi-modal Large Language Models (MLLMs) can automate the creation of accurate and coherent radiological reports. Existing methods often hallucinate details in text-b…

Causal Language ModelingHallucinationLanguage ModelingLanguage Modelling