paper-with-me

Papers

MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives

2025-01-07 · Wisdom O. Ikezogwo, Kevin Zhang, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo, Linda Shapiro, Ranjay Krishna

We propose MedicalNarratives, a dataset curated from medical pedagogical videos similar in nature to data collected in Think-Aloud studies and inspired by Localized Narratives, which collects grounded image-text data by curating instructors' speech and mouse cursor movements synchronized in time. MedicalNarratives enables pretraining of both semantic and dense objectives, alleviating the need to train medical semantic and dense tasks disparately due to the lack of reasonably sized datasets. Our dataset contains 4.7M image-text pairs from videos and articles, with 1M samples containing dense annotations in the form of traces and bounding boxes. To evaluate the utility of MedicalNarratives, we train GenMedClip based on the CLIP architecture using our dataset spanning 12 medical domains and demonstrate that it outperforms previous state-of-the-art models on a newly constructed medical imaging benchmark that comprehensively evaluates performance across all modalities. Data, demo, code and models available at https://medical-narratives.github.io

📄 PDF Abstract BibTeX arXiv:2501.04184

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Connecting Vision and Language with Video Localized Narratives

2023-02-22 · CVPR 2023 1 · Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset, Radu Soricut 외

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, th…

Question AnsweringVideo Narrative GroundingVideo Question Answering

Connecting Vision and Language with Localized Narratives

2019-12-06 · ECCV 2020 8 · Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut 외

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the regio…

FormImage CaptioningImage GenerationVisual Grounding

MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding

2025-06-10 · Shivang Chopra, Lingchao Mao, Gabriela Sanchez-Rodriguez, Andrew J Feola 외

Different medical imaging modalities capture diagnostic information at varying spatial resolutions, from coarse global patterns to fine-grained localized structures. However, most existing vision-language frameworks in t…

DiagnosticMixture-of-Experts

Joint Learning of Localized Representations from Medical Images and Reports

2021-12-06 · Philip Müller, Georgios Kaissis, Congyu Zou, Daniel Rueckert

Contrastive learning has proven effective for pre-training image models on unlabeled data with promising results for tasks such as medical image classification. Using paired text (like radiological reports) during pre-tr…

Contrastive Learningimage-classificationMedical Image Classificationobject-detection+4

Describe Anything in Medical Images

2025-05-09 · Xi Xiao, Yunbei Zhang, Thanh-Huy Nguyen, Ba-Thinh Lam 외

Localized image captioning has made significant progress with models like the Describe Anything Model (DAM), which can generate detailed region-specific descriptions without explicit region-text supervision. However, suc…

AttributeDiagnosticImage Captioning