paper-with-me

Papers

Improving Medical Report Generation with Adapter Tuning and Knowledge Enhancement in Vision-Language Foundation Models

2023-12-07 · Shibin Wu, Bang Yang, Zhiyu Ye, Haoqian Wang, Hairong Zheng, Tong Zhang

Medical report generation demands automatic creation of coherent and precise descriptions for medical images. However, the scarcity of labelled medical image-report pairs poses formidable challenges in developing large-scale neural networks capable of harnessing the potential of artificial intelligence, exemplified by large language models. This study builds upon the state-of-the-art vision-language pre-training and fine-tuning approach, BLIP-2, to customize general large-scale foundation models. Integrating adapter tuning and a medical knowledge enhancement loss, our model significantly improves accuracy and coherence. Validation on the dataset of ImageCLEFmedical 2023 demonstrates our model's prowess, achieving the best-averaged results against several state-of-the-art methods. Significant improvements in ROUGE and CIDEr underscore our method's efficacy, highlighting promising outcomes for the rapid medical-domain adaptation of the vision-language foundation models in addressing challenges posed by data scarcity.

📄 PDF Abstract BibTeX arXiv:2312.03970

Code (1)

https://openi.pcl.ac.cn/OpenMedIA/MAKEN 공식 구현

Tasks

Domain AdaptationMedical Report Generation

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation

2025-03-20 · Yaxiong Chen, Chuang Du, Chunlei Li, Jingliang Hu 외

Automated radiology report generation aims to expedite the tedious and error-prone reporting process for radiologists. While recent works have made progress, learning to align medical images and textual findings remains …

Image CaptioningTransfer Learning

Diversifying Knowledge Enhancement of Biomedical Language Models using Adapter Modules and Knowledge Graphs

2023-12-21 · Juraj Vladika, Alexander Fichtl, Florian Matthes

Recent advances in natural language processing (NLP) owe their success to pre-training language models on large amounts of unstructured data. Still, there is an increasing effort to combine the unstructured nature of LMs…

Document ClassificationKnowledge GraphsNatural Language InferenceQuestion Answering

FedPIA -- Permuting and Integrating Adapters leveraging Wasserstein Barycenters for Finetuning Foundation Models in Multi-Modal Federated Learning

2024-12-19 · Pramit Saha, Divyanshu Mishra, Felix Wagner, Konstantinos Kamnitsas 외

Large Vision-Language Models typically require large text and image datasets for effective fine-tuning. However, collecting data from various sites, especially in healthcare, is challenging due to strict privacy regulati…

Federated Learningparameter-efficient fine-tuningQuestion AnsweringVisual Question Answering

Diff-CXR: Report-to-CXR generation through a disease-knowledge enhanced diffusion model

2024-10-26 · Peng Huang, Bowen Guo, Shuyu Liang, Junhu Fu 외

Text-To-Image (TTI) generation is significant for controlled and diverse image generation with broad potential applications. Although current medical TTI methods have made some progress in report-to-Chest-Xray (CXR) gene…

Image Generation

Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation

2021-11-08 · NeurIPS 2021 12 · Fenglin Liu, Chenyu You, Xian Wu, Shen Ge 외

Medical report generation, which aims to automatically generate a long and coherent report of a given medical image, has been receiving growing research interests. Existing approaches mainly adopt a supervised manner and…

DecoderMedical Report Generation