paper-with-me

홈 › Papers

MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model

2022-10-11 · CVPR 2023 1 · Yatai Ji, Junjie Wang, Yuan Gong, Lin Zhang, Yanru Zhu, Hongfa Wang, Jiaxing Zhang, Tetsuya Sakai, Yujiu Yang

Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty, particularly in pre-training on unlabeled datasets and fine-tuning in task-specific downstream datasets. In this paper, we project the representations of all modalities as probabilistic distributions via a Probability Distribution Encoder (PDE) by utilizing sequence-level interactions. Compared to the existing deterministic methods, such uncertainty modeling can convey richer multimodal semantic information and more complex relationships. Furthermore, we integrate uncertainty modeling with popular pre-training frameworks and propose suitable pre-training tasks: Distribution-based Vision-Language Contrastive learning (D-VLC), Distribution-based Masked Language Modeling (D-MLM), and Distribution-based Image-Text Matching (D-ITM). The fine-tuned models are applied to challenging downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment, and achieve state-of-the-art results.

📄 PDF Abstract BibTeX arXiv:2210.05335

Code (1)

iigroup/map 공식 구현 pytorch

Tasks

Contrastive LearningImage-text matchingImage-text RetrievalLanguage ModelingLanguage ModellingMasked Language ModelingQuestion AnsweringRetrievalText MatchingText RetrievalVisual EntailmentVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

2026-03-22 · Jingchen Sun, Shaobo Han, Deep Patel, Wataru Kohno 외 arxiv

Knowledge distillation establishes a learning paradigm that leverages both data supervision and teacher guidance. However, determining the optimal balance between learning from data and learning from the teacher is chall…

Knowledge Distillation

Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

2026-05-06 · Huatian Zhang, Zhendong Mao, Lei Zhang, Yongdong Zhang arxiv

Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pairs. One of its key challenges lies in ho…

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

2026-03-25 · Chung-En Johnny Yu, Brian Jalaian, Nathaniel D. Bastian arxiv

Combining multiple Vision-Language Models (VLMs) can enhance multimodal reasoning and robustness, but aggregating heterogeneous models' outputs amplifies uncertainty and increases the risk of hallucinations. We propose S…

Multimodal Reasoning

Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment

2025-11-06 · Zehui Feng, Chenqi Zhang, Mingru Wang, Minuo Wei 외 arxiv

Unveiling visual semantics from neural signals such as EEG, MEG, and fMRI remains a fundamental challenge due to subject variability and the entangled nature of visual features. Existing approaches primarily align neural…

Uncertainty-Aware Vision-Language Segmentation for Medical Imaging

2026-02-16 · Aryan Das, Tanishq Rachamalla, Koushik Biswas, Swalpa Kumar Roy 외 arxiv

We introduce a novel uncertainty-aware multimodal segmentation framework that leverages both radiological images and associated clinical text for precise medical diagnosis. We propose a Modality Decoding Attention Block …

Medical Diagnosis