paper-with-me

Papers

Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language Analysis

2019-11-13 · Zhongkai Sun, Prathusha Sarma, William Sethares, YIngyu Liang

Multimodal language analysis often considers relationships between features based on text and those based on acoustical and visual properties. Text features typically outperform non-text features in sentiment analysis or emotion recognition tasks in part because the text features are derived from advanced language models or word embeddings trained on massive data sources while audio and video features are human-engineered and comparatively underdeveloped. Given that the text, audio, and video are describing the same utterance in different ways, we hypothesize that the multimodal sentiment analysis and emotion recognition can be improved by learning (hidden) correlations between features extracted from the outer product of text and audio (we call this text-based audio) and analogous text-based video. This paper proposes a novel model, the Interaction Canonical Correlation Network (ICCN), to learn such multimodal embeddings. ICCN learns correlations between all three modes via deep canonical correlation analysis (DCCA) and the proposed embeddings are then tested on several benchmark datasets and against other state-of-the-art multimodal embedding algorithms. Empirical results and ablation studies confirm the effectiveness of ICCN in capturing useful information from all three views.

📄 PDF Abstract BibTeX arXiv:1911.05544

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionMultimodal Sentiment AnalysisSentiment AnalysisWord Embeddings

Similar Papers 제목 키워드 기반

Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval

2019-08-10 · Donghuo Zeng, Yi Yu, Keizo Oyama

Cross-modal retrieval aims to retrieve data in one modality by a query in another modality, which has been a very interesting research issue in the field of multimedia, information retrieval, and computer vision, and dat…

Cross-Modal RetrievalInformation RetrievalRetrievalTriplet

Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation

2025-12-02 · Jianzong Wu, Hao Lian, Dachao Hao, Ye Tian 외 arxiv

Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: Does audio-video joint denoising training…

Video Generation

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA

2019-08-10 · Donghuo Zeng, Yi Yu, Keizo Oyama

Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal st…

audio-visual learningRetrievalVideo Retrieval

VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild

2022-11-27 · Kun Cheng, Xiaodong Cun, Yong Zhang, Menghan Xia 외

We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disen…

Video EditingVideo Generation

Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition

2024-05-26 · Tong Shi, Xuri Ge, Joemon M. Jose, Nicolas Pugeault 외

Capturing complex temporal relationships between video and audio modalities is vital for Audio-Visual Emotion Recognition (AVER). However, existing methods lack attention to local details, such as facial state changes be…

Emotion RecognitionOptical Flow Estimation