paper-with-me

홈 › Papers

GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-Efficient Medical Image Recognition

2021-01-01 · ICCV 2021 10 · Shih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena Yeung

In recent years, the growing number of medical imaging studies is placing an ever-increasing burden on radiologists. Deep learning provides a promising solution for automatic medical image analysis and clinical decision support. However, large-scale manually labeled datasets required for training deep neural networks are difficult and expensive to obtain for medical images. The purpose of this work is to develop label-efficient multimodal medical imaging representations by leveraging radiology reports. Specifically, we propose an attention-based framework (GLoRIA) for learning global and local representations by contrasting image sub-regions and words in the paired report. In addition, we propose methods to leverage the learned representations for various downstream medical image recognition tasks with limited labels. Our results demonstrate high-performance and label-efficiency for image-text retrieval, classification (finetuning and zeros-shot settings), and segmentation on different datasets.

📄 PDF Abstract BibTeX

Code (2)

marshuang80/gloria 공식 구현 pytorch
jbdel/vilmedic pytorch

Tasks

Image-text RetrievalMedical Image AnalysisRepresentation LearningRetrievalText Retrieval

Similar Papers 제목 키워드 기반

Trimodal Glioma Representation Alignment via Volumetric Contrastive Learning

2026-06-12 · Denise Marini, Eleonora Grassucci, Danilo Comminiello arxiv

Glioma grading and survival prediction require the integration of heterogeneous information collected at different spatial and biological scales. Histopathology describes tissue morphology, mRNA expression captures molec…

Contrastive Learning

GLoRIA: Gated Low-Rank Interpretable Adaptation for Dialectal ASR

2026-03-02 · Pouya Mehralian, Melissa Farasyn, Anne Breitbarth, Anne-Sophie Ghyselen 외 arxiv

Automatic Speech Recognition (ASR) in dialect-heavy settings remains challenging due to strong regional variation and limited labeled data. We propose GLoRIA, a parameter-efficient adaptation framework that leverages met…

Speech Recognition

Local-Global Multimodal Contrastive Learning for Molecular Property Prediction

2026-01-30 · Xiayu Liu, Zhengyi Lu, Yunhong Liao, Chan Fan 외 arxiv

Accurate molecular property prediction requires integrating complementary information from molecular structure and chemical semantics. In this work, we propose LGM-CL, a local-global multimodal contrastive learning frame…

Molecular Property PredictionRepresentation LearningContrastive Learning

Multimodal Federated Learning via Contrastive Representation Ensemble

2023-02-17 · Qiying Yu, Yang Liu, Yimu Wang, Ke Xu 외

With the increasing amount of multimedia data on modern mobile systems and IoT infrastructures, harnessing these rich multimodal data without breaching user privacy becomes a critical issue. Federated learning (FL) serve…

Federated LearningImage-text RetrievalQuestion AnsweringRetrieval+3

T3M: Text Guided 3D Human Motion Synthesis from Speech

2024-08-23 · Wenshuo Peng, Kaipeng Zhang, Sai Qian Zhang

Speech-driven 3D motion synthesis seeks to create lifelike animations based on human speech, with potential uses in virtual reality, gaming, and the film production. Existing approaches reply solely on speech audio for m…

DiversityMotion GenerationMotion Synthesis