Affect-DML: Context-Aware One-Shot Recognition of Human Affect using Deep Metric Learning
Human affect recognition is a well-established research area with numerous applications, e.g., in psychological care, but existing methods assume that all emotions-of-interest are given a priori as annotated training examples. However, the rising granularity and refinements of the human emotional spectrum through novel psychological theories and the increased consideration of emotions in context brings considerable pressure to data collection and labeling work. In this paper, we conceptualize one-shot recognition of emotions in context -- a new problem aimed at recognizing human affect states in finer particle level from a single support sample. To address this challenging task, we follow the deep metric learning paradigm and introduce a multi-modal emotion embedding approach which minimizes the distance of the same-emotion embeddings by leveraging complementary information of human appearance and the semantic scene context obtained through a semantic segmentation network. All streams of our context-aware model are optimized jointly using weighted triplet loss and weighted cross entropy loss. We conduct thorough experiments on both, categorical and numerical emotion recognition tasks of the Emotic dataset adapted to our one-shot recognition problem, revealing that categorizing human affect from a single example is a hard task. Still, all variants of our model clearly outperform the random baseline, while leveraging the semantic scene context consistently improves the learnt representations, setting state-of-the-art results in one-shot emotion recognition. To foster research of more universal representations of human affect states, we will make our benchmark and models publicly available to the community under https://github.com/KPeng9510/Affect-DML.
Code (1)
Tasks
Emotion RecognitionMetric LearningSemantic SegmentationTripletMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos
Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from intertwined human-human interactions, con…
Affect Recognition in Conversations Using Large Language Models
Affect recognition, encompassing emotions, moods, and feelings, plays a pivotal role in human communication. In the realm of conversational artificial intelligence, the ability to discern and respond to human affective c…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)In-Context Learningspeech-recognition+1Towards Intercultural Affect Recognition: Audio-Visual Affect Recognition in the Wild Across Six Cultures
In our multicultural world, affect-aware AI systems that support humans need the ability to perceive affect across variations in emotion expression patterns across cultures. These systems must perform well in cultural co…
Causal DiscoveryCultural Vocal Bursts Intensity Predictionfeature selectionCross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions
Robustness to domain changes is a key capability for effective deployment of human action recognition systems in real-world scenarios, where action categories at inference can present important domain shifts or even unse…
Zero-Shot Action RecognitionTransfer LearningTemporal Modeling Matters: A Novel Temporal Emotional Modeling Approach for Speech Emotion Recognition
Speech emotion recognition (SER) plays a vital role in improving the interactions between humans and machines by inferring human emotion and affective states from speech signals. Whereas recent works primarily focus on m…
Speech Emotion Recognition