paper-with-me

Papers

A scoping review on multimodal deep learning in biomedical images and texts

2023-07-14 · Zhaoyi Sun, Mingquan Lin, Qingqing Zhu, Qianqian Xie, Fei Wang, Zhiyong Lu, Yifan Peng

Computer-assisted diagnostic and prognostic systems of the future should be capable of simultaneously processing multimodal data. Multimodal deep learning (MDL), which involves the integration of multiple sources of data, such as images and text, has the potential to revolutionize the analysis and interpretation of biomedical data. However, it only caught researchers' attention recently. To this end, there is a critical need to conduct a systematic review on this topic, identify the limitations of current work, and explore future directions. In this scoping review, we aim to provide a comprehensive overview of the current state of the field and identify key concepts, types of studies, and research gaps with a focus on biomedical images and texts joint learning, mainly because these two were the most commonly available data types in MDL research. This study reviewed the current uses of multimodal deep learning on five tasks: (1) Report generation, (2) Visual question answering, (3) Cross-modal retrieval, (4) Computer-aided diagnosis, and (5) Semantic segmentation. Our results highlight the diverse applications and potential of MDL and suggest directions for future research in the field. We hope our review will facilitate the collaboration of natural language processing (NLP) and medical imaging communities and support the next generation of decision-making and computer-assisted diagnostic system development.

📄 PDF Abstract BibTeX arXiv:2307.07362

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalDecision MakingDiagnosticMultimodal Deep LearningQuestion AnsweringRetrievalSemantic SegmentationVisual Question Answering

Methods 이 논문이 사용한 방법론

Focus 설명 없음
MDL Minimum Description Length provides a criterion for the selection of models, regardless of their complexity, without the restrictive assumption that the data form a sample…

Similar Papers 제목 키워드 기반

Expediting data extraction using a large language model (LLM) and scoping review protocol: a methodological study within a complex scoping review

2025-07-09 · James Stewart-Evans, Emma Wilson, Tessa Langley, Andrew Prayle 외 arxiv

The data extraction stages of reviews are resource-intensive, and researchers may seek to expediate data extraction using online (large language models) LLMs and review protocols. Claude 3.5 Sonnet was used to trial two …

Prompt Engineering

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

2025-07-30 · Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu 외 arxiv

Large multimodal models (LMMs) have demonstrated significant potential in providing innovative solutions for various biomedical tasks, including pathology analysis, radiology report generation, and biomedical assistance.…

Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications

2024-11-06 · Daan Schouten, Giulia Nicoletti, Bas Dille, Catherine Chia 외

Recent technological advances in healthcare have led to unprecedented growth in patient data quantity and diversity. While artificial intelligence (AI) models have shown promising results in analyzing individual data mod…

Decision Making

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

2025-02-13 · Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier 외

Generative artificial intelligence (AI) models, such as diffusion models and OpenAI's ChatGPT, are transforming medicine by enhancing diagnostic accuracy and automating clinical workflows. The field has advanced rapidly,…

DiagnosticDrug DiscoveryMedical Report Generation

Trust in Human-AI Interaction: Scoping Out Models, Measures, and Methods

2022-04-30 · Takane Ueno, Yuto Sawa, Yeongdae Kim, Jacqueline Urakami 외

Trust has emerged as a key factor in people's interactions with AI-infused systems. Yet, little is known about what models of trust have been used and for what systems: robots, virtual characters, smart vehicles, decisio…