paper-with-me

Papers

InterMulti:Multi-view Multimodal Interactions with Text-dominated Hierarchical High-order Fusion for Emotion Analysis

2022-12-20 · Feng Qiu, Wanzeng Kong, Yu Ding

Humans are sophisticated at reading interlocutors' emotions from multimodal signals, such as speech contents, voice tones and facial expressions. However, machines might struggle to understand various emotions due to the difficulty of effectively decoding emotions from the complex interactions between multimodal signals. In this paper, we propose a multimodal emotion analysis framework, InterMulti, to capture complex multimodal interactions from different views and identify emotions from multimodal signals. Our proposed framework decomposes signals of different modalities into three kinds of multimodal interaction representations, including a modality-full interaction representation, a modality-shared interaction representation, and three modality-specific interaction representations. Additionally, to balance the contribution of different modalities and learn a more informative latent interaction representation, we developed a novel Text-dominated Hierarchical High-order Fusion(THHF) module. THHF module reasonably integrates the above three kinds of representations into a comprehensive multimodal interaction representation. Extensive experimental results on widely used datasets, (i.e.) MOSEI, MOSI and IEMOCAP, demonstrate that our method outperforms the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2212.10030

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognitionmultimodal interaction

Similar Papers 제목 키워드 기반

S5 Framework: A Review of Self-Supervised Shared Semantic Space Optimization for Multimodal Zero-Shot Learning

2022-01-16 · ACL ARR January 2022 1 · Anonymous

In this review, we aim to inspire research into Self-Supervised Shared Semantic Space (S5) multimodal learning problems. We equip non-expert researchers with a framework of informed modeling decisions via an extensive li…

DenoisingZero-Shot Learning

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

2025-09-04 · Xiyuan Gao, Shekhar Nayak, Matt Coler arxiv

Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in …

Sarcasm Detection

Towards Multimodal Understanding of Passenger-Vehicle Interactions in Autonomous Vehicles: Intent/Slot Recognition Utilizing Audio-Visual Data

2019-09-20 · Eda Okur, Shachi H. Kumar, Saurav Sahay, Lama Nachman

Understanding passenger intents from spoken interactions and car's vision (both inside and outside the vehicle) are important building blocks towards developing contextual dialog systems for natural interactions in auton…

Autonomous VehiclesIntent Detectionslot-fillingSlot Filling

cPAPERS: A Dataset of Situated and Multimodal Interactive Conversations in Scientific Papers

2024-06-12 · Anirudh Sundar, Jin Xu, William Gay, Christopher Richardson 외

An emerging area of research in situated and multimodal interactive conversations (SIMMC) includes interactions in scientific papers. Since scientific papers are primarily composed of text, equations, figures, and tables…

MCiteBench: A Multimodal Benchmark for Generating Text with Citations

2025-03-04 · Caiyu Hu, Yikai Zhang, Tinghui Zhu, Yiwei Ye 외

Multimodal Large Language Models (MLLMs) have advanced in integrating diverse modalities but frequently suffer from hallucination. A promising solution to mitigate this issue is to generate text with citations, providing…

HallucinationText Generation