paper-with-me

홈 › Papers

Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training

2023-05-13 · Ke Zhang, Yan Yang, Jun Yu, Hanliang Jiang, Jianping Fan, Qingming Huang, Weidong Han

In recent years, the growing demand for medical imaging diagnosis has placed a significant burden on radiologists. As a solution, Medical Vision-Language Pre-training (Med-VLP) methods have been proposed to learn universal representations from medical images and reports, benefiting downstream tasks without requiring fine-grained annotations. However, existing methods have overlooked the importance of cross-modal alignment in joint image-text reconstruction, resulting in insufficient cross-modal interaction. To address this limitation, we propose a unified Med-VLP framework based on Multi-task Paired Masking with Alignment (MPMA) to integrate the cross-modal alignment task into the joint image-text reconstruction framework to achieve more comprehensive cross-modal interaction, while a Global and Local Alignment (GLA) module is designed to assist self-supervised paradigm in obtaining semantic representations with rich domain knowledge. Furthermore, we introduce a Memory-Augmented Cross-Modal Fusion (MA-CMF) module to fully integrate visual information to assist report reconstruction and fuse the multi-modal representations adequately. Experimental results demonstrate that the proposed unified approach outperforms previous methods in all downstream tasks, including uni-modal, cross-modal, and multi-modal tasks.

📄 PDF Abstract BibTeX arXiv:2305.07920

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignment

Similar Papers 제목 키워드 기반

SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining

2024-04-01 · CVPR 2024 1 · Chull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi 외

Vision-language models (VLMs) have made significant strides in cross-modal understanding through large-scale paired datasets. However, in fashion domain, datasets often exhibit a disparity between the information conveye…

Contrastive LearningImage-text matchingLanguage ModelingLanguage Modelling+2

Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations

2025-03-25 · CVPR 2025 1 · Jungin Park, Jiyoung Lee, Kwanghoon Sohn

View-invariant representation learning from egocentric (first-person, ego) and exocentric (third-person, exo) videos is a promising approach toward generalizing video understanding systems across multiple viewpoints. How…

Representation LearningVideo Understanding

MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling

2021-09-24 · Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman, Yi Xu 외

Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (especially image embedders), (ii) loss f…

Image ReconstructionLanguage ModelingLanguage ModellingMasked Language Modeling

UNITER: UNiversal Image-TExt Representation Learning

2019-09-25 · ECCV 2020 8 · Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy 외

Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, …

Image-text matchingImage-text RetrievalLanguage ModelingLanguage Modelling+14

Data-efficient Event Camera Pre-training via Disentangled Masked Modeling

2024-03-01 · Zhenpeng Huang, Chao Li, Hao Chen, Yongjian Deng 외

In this paper, we present a new data-efficient voxel-based self-supervised learning method for event cameras. Our pre-training overcomes the limitations of previous methods, which either sacrifice temporal information by…

Knowledge DistillationSelf-Supervised Learning