paper-with-me

Papers

Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis

2026-07-14 · Alexios Filippakopoulos, Elias Kallioras, Nikolaos Xiros, Efthymios Georgiou, Alexandros Potamianos arxiv

Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.

📄 PDF Abstract BibTeX arXiv:2607.12686

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Similar Papers 제목 키워드 기반

NoteLLM-2: Multimodal Large Representation Models for Recommendation

2024-05-27 · Chao Zhang, Haoxin Zhang, Shiwei Wu, Di wu 외

Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularly for item-to-item (I2I) recommendations…

In-Context Learning

Attention-Based Multimodal Survival Prediction with Cross-Modal Bilinear Fusion

2026-05-12 · Hassan Keshvarikhojasteh, Josien P. W. Pluim, Mitko Veta arxiv

We propose a novel multimodal deep learning framework for patient-level survival prediction, which integrates whole-slide histology features, RNA-seq expression profiles, and clinical variables. Our architecture combines…

Multimodal Deep LearningMultimodal Reasoning

Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion

2021-11-01 · EMNLP 2021 11 · Xiaobao Guo, Adams Kong, Huan Zhou, Xianfeng Wang 외

Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning. Prior works often modulate one modal feature to another straightforwardly and…

Representation Learning

FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction

2024-06-17 · Muhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng 외

Multimodal electronic health record (EHR) data can offer a holistic assessment of a patient's health status, supporting various predictive healthcare tasks. Recently, several studies have embraced the multitask learning …

FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking

2025-10-22 · Martha Teiko Teye, Ori Maoz, Matthias Rottmann arxiv

We propose FutrTrack, a modular camera-LiDAR multi-object tracking framework that builds on existing 3D detectors by introducing a transformer-based smoother and a fusion-driven tracker. Inspired by query-based tracking …

Multiple Object TrackingMulti-Object Tracking