paper-with-me

홈 › Papers

Pairwise Emotional Relationship Recognition in Drama Videos: Dataset and Benchmark

2021-09-23 · Xun Gao, Yin Zhao, Jie Zhang, Longjun Cai

Recognizing the emotional state of people is a basic but challenging task in video understanding. In this paper, we propose a new task in this field, named Pairwise Emotional Relationship Recognition (PERR). This task aims to recognize the emotional relationship between the two interactive characters in a given video clip. It is different from the traditional emotion and social relation recognition task. Varieties of information, consisting of character appearance, behaviors, facial emotions, dialogues, background music as well as subtitles contribute differently to the final results, which makes the task more challenging but meaningful in developing more advanced multi-modal models. To facilitate the task, we develop a new dataset called Emotional RelAtionship of inTeractiOn (ERATO) based on dramas and movies. ERATO is a large-scale multi-modal dataset for PERR task, which has 31,182 video clips, lasting about 203 video hours. Different from the existing datasets, ERATO contains interaction-centric videos with multi-shots, varied video length, and multiple modalities including visual, audio and text. As a minor contribution, we propose a baseline model composed of Synchronous Modal-Temporal Attention (SMTA) unit to fuse the multi-modal information for the PERR task. In contrast to other prevailing attention mechanisms, our proposed SMTA can steadily improve the performance by about 1\%. We expect the ERATO as well as our proposed SMTA to open up a new way for PERR task in video understanding and further improve the research of multi-modal fusion methodology.

📄 PDF Abstract BibTeX arXiv:2109.11243

Code (1)

cti-vision/perr 공식 구현

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios

2024-04-11 · Yuan Zhang, Xiaomei Tao, Hanxu Ai, Tao Chen 외

In the Massive Open Online Courses (MOOC) learning scenario, the semantic information of instructional videos has a crucial impact on learners' emotional state. Learners mainly acquire knowledge by watching instructional…

Emotion RecognitionLanguage ModellingLarge Language ModelMultimodal Emotion Recognition+1

DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation

2025-12-22 · Shijian Ma, Yunqi Huang, Yan Lin arxiv

Drama script continuation requires models to maintain character consistency, advance plot coherently, and preserve dramatic structurecapabilities that existing benchmarks fail to evaluate comprehensively. We present Dram…

Action Genome: Actions as Composition of Spatio-temporal Scene Graphs

2019-12-15 · Jingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos Niebles

Action recognition has typically treated actions and activities as monolithic events that occur in videos. However, there is evidence from Cognitive Science and Neuroscience that people actively encode activities into co…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionSpatio-temporal Scene Graphs

Action Genome: Actions As Compositions of Spatio-Temporal Scene Graphs

2020-06-01 · CVPR 2020 6 · Jingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos Niebles

Action recognition has typically treated actions and activities as monolithic events that occur in videos. However, there is evidence from Cognitive Science and Neuroscience that people actively encode activities into co…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionSpatio-temporal Scene Graphs

Taming Transformer for Emotion-Controllable Talking Face Generation

2025-08-20 · Ziqi Zhang, Cheng Deng arxiv

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need…

Talking Face Generation