paper-with-me

Papers

Asymmetric Reinforcing against Multi-modal Representation Bias

2025-01-02 · Xiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang, QinGhua Hu

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning.

📄 PDF Abstract BibTeX arXiv:2501.01240

Code (1)

gao-xiyuan/arm 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Reinforcing Multimodal Reasoning Against Visual Degradation

2026-05-10 · Rui Liu, Dian Yu, Haolin Liu, Yucheng Shi 외 arxiv

Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies remain brittle against real-world visual degradations such as blur, com…

Reinforcement LearningMultimodal ReasoningData Augmentation

Protein Representation Learning by Capturing Protein Sequence-Structure-Function Relationship

2024-04-29 · Eunji Ko, Seul Lee, Minseon Kim, DongKi Kim

The goal of protein representation learning is to extract knowledge from protein databases that can be applied to various protein-related downstream tasks. Although protein sequence, structure, and function are the three…

Representation Learning

Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search

2025-06-25 · Zhigong Zhou, Ning Ding, Xiaochuan Fan, Yue Shang 외

Semantic retrieval, which retrieves semantically matched items given a textual query, has been an essential component to enhance system effectiveness in e-commerce search. In this paper, we study the multimodal retrieval…

Question AnsweringRetrievalSemantic RetrievalVisual Question Answering

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

2026-05-21 · Changyuan Tian, Zhicong Lu, Huaxing Liu, Xiang Wang 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to multimodal large language models (MLLMs)…

Reinforcement LearningMultimodal Reasoning

AsymFormer: Asymmetrical Cross-Modal Representation Learning for Mobile Platform Real-Time RGB-D Semantic Segmentation

2023-09-25 · Siqi Du, Weixi Wang, Renzhong Guo, Ruisheng Wang 외

Understanding indoor scenes is crucial for urban studies. Considering the dynamic nature of indoor environments, effective semantic segmentation requires both real-time operation and high accuracy.To address this, we pro…

Computational EfficiencyFeature Correlationfeature selectionQuantization+4