Medical visual question answering using joint self-supervised learning
Visual Question Answering (VQA) becomes one of the most active research problems in the medical imaging domain. A well-known VQA challenge is the intrinsic diversity between the image and text modalities, and in the medical VQA task, there is another critical problem relying on the limited size of labelled image-question-answer data. In this study we propose an encoder-decoder framework that leverages the image-text joint representation learned from large-scaled medical image-caption data and adapted to the small-sized medical VQA task. The encoder embeds across the image-text dual modalities with self-attention mechanism and is independently pre-trained on the large-scaled medical image-caption dataset by multiple self-supervised learning tasks. Then the decoder is connected to the top of the encoder and fine-tuned using the small-sized medical VQA dataset. The experiment results present that our proposed method achieves better performance comparing with the baseline and SOTA methods.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderDiversityMedical Visual Question AnsweringQuestion AnsweringSelf-Supervised LearningVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Towards Visual Question Answering on Pathology Images
Pathology imaging is broadly used for identifying the causes and effects of diseases or injuries. Given a pathology image, being able to answer questions about the clinical findings contained in the image is very importa…
Decision MakingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Self-supervised vision-language pretraining for Medical visual question answering
Medical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information. To…
Contrastive LearningImage-text matchingLanguage ModelingLanguage Modelling+6Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
Medical Visual Question Answering (Med-VQA) answers clinical questions using medical images, aiding diagnosis. Designing the MedVQA system holds profound importance in assisting clinical diagnosis and enhancing diagnosti…
DiagnosticMedical Visual Question AnsweringQuestion AnsweringVisual Question Answering+1STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medical image understanding and reasoning crit…
Medical DiagnosisMedical Question AnsweringMedical Visual Question AnsweringQuestion Answering+2Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering
Medical visual question answering (VQA) is a challenging task that requires answering clinical questions of a given medical image, by taking consider of both visual and language information. However, due to the small sca…
Language ModelingMedical Visual Question Answering