paper-with-me

Papers

Medical visual question answering using joint self-supervised learning

2023-02-25 · Yuan Zhou, Jing Mei, Yiqin Yu, Tanveer Syeda-Mahmood

Visual Question Answering (VQA) becomes one of the most active research problems in the medical imaging domain. A well-known VQA challenge is the intrinsic diversity between the image and text modalities, and in the medical VQA task, there is another critical problem relying on the limited size of labelled image-question-answer data. In this study we propose an encoder-decoder framework that leverages the image-text joint representation learned from large-scaled medical image-caption data and adapted to the small-sized medical VQA task. The encoder embeds across the image-text dual modalities with self-attention mechanism and is independently pre-trained on the large-scaled medical image-caption dataset by multiple self-supervised learning tasks. Then the decoder is connected to the top of the encoder and fine-tuned using the small-sized medical VQA dataset. The experiment results present that our proposed method achieves better performance comparing with the baseline and SOTA methods.

📄 PDF Abstract BibTeX arXiv:2302.13069

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDiversityMedical Visual Question AnsweringQuestion AnsweringSelf-Supervised LearningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Towards Visual Question Answering on Pathology Images

2021-08-01 · ACL 2021 5 · Xuehai He, Zhuo Cai, Wenlan Wei, Yichen Zhang 외

Pathology imaging is broadly used for identifying the causes and effects of diseases or injuries. Given a pathology image, being able to answer questions about the clinical findings contained in the image is very importa…

Decision MakingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Self-supervised vision-language pretraining for Medical visual question answering

2022-11-24 · Pengfei Li, Gang Liu, Lin Tan, Jinying Liao 외

Medical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information. To…

Contrastive LearningImage-text matchingLanguage ModelingLanguage Modelling+6

Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion

2025-04-04 · Junkai Zhang, Bin Li, Shoujun Zhou, Yue Du

Medical Visual Question Answering (Med-VQA) answers clinical questions using medical images, aiding diagnosis. Designing the MedVQA system holds profound importance in assisting clinical diagnosis and enhancing diagnosti…

DiagnosticMedical Visual Question AnsweringQuestion AnsweringVisual Question Answering+1

STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering

2024-06-28 · Guohao Sun, Can Qin, Huazhu Fu, Linwei Wang 외

Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medical image understanding and reasoning crit…

Medical DiagnosisMedical Question AnsweringMedical Visual Question AnsweringQuestion Answering+2

Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering

2023-07-11 · Pengfei Li, Gang Liu, Jinlong He, Zixu Zhao 외

Medical visual question answering (VQA) is a challenging task that requires answering clinical questions of a given medical image, by taking consider of both visual and language information. However, due to the small sca…

Language ModelingMedical Visual Question Answering