Visual-Texual Emotion Analysis with Deep Coupled Video and Danmu Neural Networks
User emotion analysis toward videos is to automatically recognize the general emotional status of viewers from the multimedia content embedded in the online video stream. Existing works fall in two categories: 1) visual-based methods, which focus on visual content and extract a specific set of features of videos. However, it is generally hard to learn a mapping function from low-level video pixels to high-level emotion space due to great intra-class variance. 2) textual-based methods, which focus on the investigation of user-generated comments associated with videos. The learned word representations by traditional linguistic approaches typically lack emotion information and the global comments usually reflect viewers' high-level understandings rather than instantaneous emotions. To address these limitations, in this paper, we propose to jointly utilize video content and user-generated texts simultaneously for emotion analysis. In particular, we introduce exploiting a new type of user-generated texts, i.e., "danmu", which are real-time comments floating on the video and contain rich information to convey viewers' emotional opinions. To enhance the emotion discriminativeness of words in textual feature extraction, we propose Emotional Word Embedding (EWE) to learn text representations by jointly considering their semantics and emotions. Afterwards, we propose a novel visual-textual emotion analysis model with Deep Coupled Video and Danmu Neural networks (DCVDN), in which visual and textual features are synchronously extracted and fused to form a comprehensive representation by deep-canonically-correlated-autoencoder-based multi-view learning. Through extensive experiments on a self-crawled real-world video-danmu dataset, we prove that DCVDN significantly outperforms the state-of-the-art baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionMULTI-VIEW LEARNINGSimilar Papers 제목 키워드 기반
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made prog…
Video GenerationUnlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion
The emergence of artificial emotional intelligence technology is revolutionizing the fields of computers and robotics, allowing for a new level of communication and understanding of human behavior that was once thought i…
Emotional IntelligenceEmotion RecognitionX-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that e…
DialogueEIN: Emotional Interaction Network for Emotion Recognition in Conversations
Emotion Recognition in Conversations (ERC) is a necessary step for developing empathetic human-computer interaction system. The existing methods on ERC primarily focus on capturing the context-level and speaker-level inf…
Emotion RecognitionEmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framew…
Video Generation