paper-with-me

홈 › Papers

Enhancing Multimodal Unified Representations for Cross Modal Generalization

2024-03-08 · Hai Huang, Yan Xia, Shengpeng Ji, Shulei Wang, Hanting Wang, Minghui Fang, Jieming Zhu, Zhenhua Dong, Sashuai Zhou, Zhou Zhao

To enhance the interpretability of multimodal unified representations, many studies have focused on discrete unified representations. These efforts typically start with contrastive learning and gradually extend to the disentanglement of modal information, achieving solid multimodal discrete unified representations. However, existing research often overlooks two critical issues: 1) The use of Euclidean distance for quantization in discrete representations often overlooks the important distinctions among different dimensions of features, resulting in redundant representations after quantization; 2) Different modalities have unique characteristics, and a uniform alignment approach does not fully exploit these traits. To address these issues, we propose Training-free Optimization of Codebook (TOC) and Fine and Coarse cross-modal Information Disentangling (FCID). These methods refine the unified discrete representations from pretraining and perform fine- and coarse-grained information disentanglement tailored to the specific characteristics of each modality, achieving significant performance improvements over previous state-of-the-art models.

📄 PDF Abstract BibTeX arXiv:2403.05168

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDisentanglementQuantizationRepresentation Learning

Similar Papers 제목 키워드 기반

UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning

2025-03-27 · Hongxuan Tang, Hao liu, Xinyan Xiao

We introduce UGen, a unified autoregressive multimodal model that demonstrates strong performance across text processing, image understanding, and image generation tasks simultaneously. UGen converts both texts and image…

Image Generation

Semantic Residual for Multimodal Unified Discrete Representation

2024-12-26 · Hai Huang, Shulei Wang, Yan Xia

Recent research in the domain of multimodal unified representations predominantly employs codebook as representation forms, utilizing Vector Quantization(VQ) for quantization, yet there has been insufficient exploration …

DisentanglementQuantizationRetrieval

Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

2025-07-04 · Hai Huang, Yan Xia, Sashuai Zhou, Hanting Wang 외

Domain Generalization (DG) aims to enhance model robustness in unseen or distributionally shifted target domains through training exclusively on source domains. Although existing DG techniques, such as data manipulation,…

DisentanglementDomain Generalization

Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention

2024-10-19 · Yuzhe Weng, Haotian Wang, Tian Gao, Kewei Li 외

In multimodal sentiment analysis, collecting text data is often more challenging than video or audio due to higher annotation costs and inconsistent automatic speech recognition (ASR) quality. To address this challenge, …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multimodal Sentiment AnalysisSentiment Analysis+2

Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning

2024-02-06 · Trilok Padhi, Ugur Kursuncu, Yaman Kumar, Valerie L. Shalin 외

The digital landscape continually evolves with multimodality, enriching the online experience for users. Creators and marketers aim to weave subtle contextual cues from various modalities into congruent content to engage…

Common Sense ReasoningKnowledge GraphsMarketing