paper-with-me

Papers

Self-supervised vision-language pretraining for Medical visual question answering

2022-11-24 · Pengfei Li, Gang Liu, Lin Tan, Jinying Liao, Shenjun Zhong

Medical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information. To solve medical VQA problems with a limited number of training data, pretrain-finetune paradigm is widely used to improve the model generalization. In this paper, we propose a self-supervised method that applies Masked image modeling, Masked language modeling, Image text matching and Image text alignment via contrastive learning (M2I2) for pretraining on medical image caption dataset, and finetunes to downstream medical VQA tasks. The proposed method achieves state-of-the-art performance on all the three public medical VQA datasets. Our codes and models are available at https://github.com/pengfeiliHEU/M2I2.

📄 PDF Abstract BibTeX arXiv:2211.13594

Code (2)

pengfeiliheu/m2i2 공식 구현 pytorch
pengfeiliheu/mumc pytorch

Tasks

Contrastive LearningImage-text matchingLanguage ModelingLanguage ModellingMasked Language ModelingMedical Visual Question AnsweringQuestion AnsweringText MatchingVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision

2026-03-14 · Yunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu 외 arxiv

Foundation models have transformed vision and language by learning general-purpose representations from large-scale unlabeled data, yet 3D medical imaging lacks analogous approaches. Existing self-supervised methods rely…

Self-Supervised Learning

Medical Vision Language Pretraining: A survey

2023-12-11 · Prashant Shrestha, Sanskar Amgain, Bidur Khanal, Cristian A. Linte 외

Medical Vision Language Pretraining (VLP) has recently emerged as a promising solution to the scarcity of labeled data in the medical domain. By leveraging paired/unpaired vision and text datasets through self-supervised…

Self-Supervised LearningSurvey

EndoViT: pretraining vision transformers on a large collection of endoscopic images

2024-04-03 · International Journal of Computer Assisted Radiology and Surgery 19:1085–109 2024 4 · Dominik Bati´c, Felix Holm, Ege Özsoy, Tobias Czempiel 외

Automated endoscopy video analysis is essential for assisting surgeons during medical procedures, but it faces challenges due to complex surgical scenes and limited annotated data. Large-scale pretraining has shown great…

Action Triplet RecognitionSegmentationSemantic SegmentationTriplet

Self-Supervised Pretraining for 2D Medical Image Segmentation

2022-09-01 · András Kalapos, Bálint Gyires-Tóth

Supervised machine learning provides state-of-the-art solutions to a wide range of computer vision problems. However, the need for copious labelled training data limits the capabilities of these algorithms in scenarios w…

Cardiac SegmentationImage SegmentationMedical Image SegmentationSegmentation+2

Exploring the Utility of Self-Supervised Pretraining Strategies for the Detection of Absent Lung Sliding in M-Mode Lung Ultrasound

2023-04-05 · Blake VanBerlo, Brian Li, Alexander Wong, Jesse Hoey 외

Self-supervised pretraining has been observed to improve performance in supervised learning tasks in medical imaging. This study investigates the utility of self-supervised pretraining prior to conducting supervised fine…

Data Augmentation