paper-with-me

Papers

A Survey of Medical Vision-and-Language Applications and Their Techniques

2024-11-19 · Qi Chen, Ruoshan Zhao, Sinuo Wang, Vu Minh Hieu Phan, Anton Van Den Hengel, Johan Verjans, Zhibin Liao, Minh-Son To, Yong Xia, Jian Chen, Yutong Xie, Qi Wu

Medical vision-and-language models (MVLMs) have attracted substantial interest due to their capability to offer a natural language interface for interpreting complex medical data. Their applications are versatile and have the potential to improve diagnostic accuracy and decision-making for individual patients while also contributing to enhanced public health monitoring, disease surveillance, and policy-making through more efficient analysis of large data sets. MVLMS integrate natural language processing with medical images to enable a more comprehensive and contextual understanding of medical images alongside their corresponding textual information. Unlike general vision-and-language models trained on diverse, non-specialized datasets, MVLMs are purpose-built for the medical domain, automatically extracting and interpreting critical information from medical images and textual reports to support clinical decision-making. Popular clinical applications of MVLMs include automated medical report generation, medical visual question answering, medical multimodal segmentation, diagnosis and prognosis and medical image-text retrieval. Here, we provide a comprehensive overview of MVLMs and the various medical tasks to which they have been applied. We conduct a detailed analysis of various vision-and-language model architectures, focusing on their distinct strategies for cross-modal integration/exploitation of medical visual and textual features. We also examine the datasets used for these tasks and compare the performance of different models based on standardized evaluation metrics. Furthermore, we highlight potential challenges and summarize future research trends and directions. The full collection of papers and codes is available at: https://github.com/YtongXie/Medical-Vision-and-Language-Tasks-and-Methodologies-A-Survey.

📄 PDF Abstract BibTeX arXiv:2411.12195

Code (1)

ytongxie/medical-vision-and-language-tasks-and-methodologies-a-survey 공식 구현 pytorch

Tasks

Decision MakingDiagnosticImage-text RetrievalMedical Report GenerationMedical Visual Question AnsweringPrognosisQuestion AnsweringText RetrievalVisual Question Answering

Similar Papers 제목 키워드 기반

Transformers in Medical Imaging: A Survey

2022-01-24 · Fahad Shamshad, Salman Khan, Syed Waqas Zamir, Muhammad Haris Khan 외

Following unprecedented success on the natural language tasks, Transformers have been successfully applied to several computer vision problems, achieving state-of-the-art results and prompting researchers to reconsider t…

Image ClassificationImage SegmentationMedical Image DenoisingMedical Image Registration+5

A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

2024-05-14 · Hanguang Xiao, Feizhong Zhou, Xingyue Liu, Tianqi Liu 외

Since the release of ChatGPT and GPT-4, large language models (LLMs) and multimodal large language models (MLLMs) have attracted widespread attention for their exceptional capabilities in understanding, reasoning, and ge…

Survey

Pre-trained Language Models in Biomedical Domain: A Systematic Survey

2021-10-11 · Benyou Wang, Qianqian Xie, Jiahuan Pei, Zhihong Chen 외

Pre-trained language models (PLMs) have been the de facto paradigm for most natural language processing (NLP) tasks. This also benefits biomedical domain: researchers from informatics, medicine, and computer science (CS)…

Survey

A Recent Survey of Vision Transformers for Medical Image Segmentation

2023-12-01 · Asifullah Khan, Zunaira Rauf, Abdul Rehman Khan, Saima Rathore 외

Medical image segmentation plays a crucial role in various healthcare applications, enabling accurate diagnosis, treatment planning, and disease monitoring. Traditionally, convolutional neural networks (CNNs) dominated t…

Image SegmentationInductive BiasMedical Image SegmentationSegmentation+2

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

2025-07-30 · Weide Liu, Wei Zhou, Jun Liu, Ping Hu 외 arxiv

Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matc…

Medical Image Registration3D ReconstructionImage RetrievalImage Matching