paper-with-me

Papers

Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning

2023-03-10 · CVPR 2023 1 · Qian Jiang, Changyou Chen, Han Zhao, Liqun Chen, Qing Ping, Son Dinh Tran, Yi Xu, Belinda Zeng, Trishul Chilimbi

Contrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly match each other in the latent space. Yet it remains an open question how the modality alignment affects the downstream task performance. In this paper, based on an information-theoretic argument, we first prove that exact modality alignment is sub-optimal in general for downstream prediction tasks. Hence we advocate that the key of better performance lies in meaningful latent modality structures instead of perfect modality alignment. To this end, we propose three general approaches to construct latent modality structures. Specifically, we design 1) a deep feature separation loss for intra-modality regularization; 2) a Brownian-bridge loss for inter-modality regularization; and 3) a geometric consistency loss for both intra- and inter-modality regularization. Extensive experiments are conducted on two popular multi-modal representation learning frameworks: the CLIP-based two-tower model and the ALBEF-based fusion model. We test our model on a variety of tasks including zero/few-shot image classification, image-text retrieval, visual question answering, visual reasoning, and visual entailment. Our method achieves consistent improvements over existing methods, demonstrating the effectiveness and generalizability of our proposed approach on latent modality structure regularization.

📄 PDF Abstract BibTeX arXiv:2303.05952

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Image Classificationimage-classificationImage ClassificationImage-text RetrievalOpen-Ended Question AnsweringQuestion AnsweringRepresentation LearningRetrievalText RetrievalVisual EntailmentVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Multiscale Structure-Guided Latent Diffusion for Multimodal MRI Translation

2026-03-13 · Jianqiang Lin, Zhiqiang Shen, Peng Cao, Jinzhu Yang 외 arxiv

Although diffusion models have achieved remarkable progress in multi-modal magnetic resonance imaging (MRI) translation tasks, existing methods still tend to suffer from anatomical inconsistencies or degraded texture det…

Must: Maximizing Latent Capacity of Spatial Transcriptomics Data

2024-01-15 · Zelin Zang, Liangyu Li, Yongjie Xu, Chenrui Duan 외

Spatial transcriptomics (ST) technologies have revolutionized the study of gene expression patterns in tissues by providing multimodality data in transcriptomic, spatial, and morphological, offering opportunities for und…

Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI

2026-04-06 · Mingjie Li, Edward Kim, Yue Zhao, Ehsan Adeli 외 arxiv

Learning a robust Variational Autoencoder (VAE) is a fundamental step for many deep learning applications in medical image analysis, such as MRI synthesizes. Existing brain VAEs predominantly focus on single-modality dat…

Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

2024-12-12 · Zhisheng Zhong, Chengyao Wang, Yuqi Liu, Senqiao Yang 외

As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently exp…

EgoSchemaMMEMM-Vet+5

M3TR: Multi-modal Multi-label Recognition with Transformer

2021-10-01 · ACM MM 2021 10 · Jiawei Zhao, Yifan Zhao, Jia Li

Multi-label image recognition aims to recognize multiple objects simultaneously in one image. Recent ideas to solve this problem have focused on learning dependencies of label co-occurrences to enhance the high-level sem…

Multi-Label ClassificationMulti-Label Image Recognition