paper-with-me

CREMA-D

홈페이지 · 논문 28편

CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors. These clips were from 48 male and 43 female actors between the ages of 20 and 74 coming from a variety of races and ethnicities (African America, Asian, Caucasian, Hispanic, and Unspecified). Actors spoke from a selection of 12 sentences. The sentences were presented using one of six different emotions (Anger, Disgust, Fear, Happy, Neutral, and Sad) and four different emotion levels (Low, Medium, High, and Unspecified). Participants rated the emotion and emotion levels based on the combined audiovisual presentation, the video alone, and the audio alone. Due to the large number of ratings needed, this effort was crowd-sourced and a total of 2443 participants each rated 90 unique clips, 30 audio, 30 visual, and 30 audio-visual. 95% of the clips have more than 7 ratings.

Audio

벤치마크

Speech Emotion Recognition on CREMA-D 결과 19개
Talking Face Generation on CREMA-D 결과 11개
Audio Classification on CREMA-D 결과 6개
Facial Expression Recognition (FER) on CREMA-D 결과 6개
Few-Shot Audio Classification on CREMA-D 결과 6개
Video Emotion Recognition on CREMA-D 결과 4개
Self-Supervised Learning on CREMA-D 결과 1개