paper-with-me

홈 › Papers

Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation

2025-12-03 · Xiaosen Lyu, Jiayu Xiong, Yuren Chen, Wanlong Wang, Xiaoqing Dai, Jing Wang arxiv

Multimodal Emotion Recognition in Conversation (MERC) aims to predict speakers' emotions by integrating textual, acoustic, and visual cues. Existing approaches either struggle to capture complex cross-modal interactions or experience gradient conflicts and unstable training when using deeper architectures. To address these issues, we propose Cross-Space Synergy (CSS), which couples a representation component with an optimization component. Synergistic Polynomial Fusion (SPF) serves the representation role, leveraging low-rank tensor factorization to efficiently capture high-order cross-modal interactions. Pareto Gradient Modulator (PGM) serves the optimization role, steering updates along Pareto-optimal directions across competing objectives to alleviate gradient conflicts and improve stability. Experiments show that CSS outperforms existing representative methods on IEMOCAP and MELD in both accuracy and training stability, demonstrating its effectiveness in complex multimodal scenarios.

📄 PDF Abstract BibTeX arXiv:2512.03521

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Emotion Recognition

Similar Papers 제목 키워드 기반

Learning Representation and Synergy Invariances: A Povable Framework for Generalized Multimodal Face Anti-Spoofing

2025-11-18 · Xun Lin, Shuai Wang, Yi Yu, Zitong Yu 외 arxiv

Multimodal Face Anti-Spoofing (FAS) methods, which integrate multiple visual modalities, often suffer even more severe performance degradation than unimodal FAS when deployed in unseen domains. This is mainly due to two …

Face Anti-Spoofing

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

2026-08-05 · Junlin Han, Shengbang Tong, David Fan, Minghao Chen 외 hf

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities int…

Modeling Cross-vision Synergy for Unified Large Vision Model

2026-03-03 · Shengqiong Wu, Lanhu Wu, Mingyang Bao, Wenhao Xu 외 arxiv

Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, videos, and 3D data. However, existing unified LVMs primarily pursue fun…

Knowledge DistillationVisual Reasoning

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

2025-05-23 · Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang 외

This paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at simultaneously reinforcing generation and u…

Image Generationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

UMSS: Towards Unsupervised Multi-modal Semantic Segmentation

2026-07-14 · Haitian Zhang, Thai Duy Nguyen, Xiangyuan Wang, Mohan Liu 외 arxiv

Multimodal semantic segmentation (MSS) is essential for robust perception in complex environments, yet its potential remains largely untapped because of the prohibitive cost of human annotations. While unsupervised seman…

Unsupervised Semantic Segmentation