Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech production, which also encompasses speaker traits, including emotional states. To bridge this gap, we conducted a series of experiments to investigate the potential of the phonation excitation information and articulatory kinematics for SER. Due to the scarcity of training data for this purpose, we introduce a portrayed emotional dataset, STEM-E2VA, which includes audio and physiological data such as electroglottography (EGG) and electromagnetic articulography (EMA). EGG and EMA provide information of phonation excitation and articulatory kinematics, respectively. Additionally, we performed emotion recognition using estimated physiological data derived through inversion methods from speech, instead of collected EGG and EMA, to explore the feasibility of applying such physiological information in real-world SER. Experimental results confirm the effectiveness of incorporating physiological information about speech production for SER and demonstrate its potential for practical use in real-world scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech Emotion RecognitionSimilar Papers 제목 키워드 기반
DNN-HMM based Speaker Adaptive Emotion Recognition using Proposed Epoch and MFCC Features
Speech is produced when time varying vocal tract system is excited with time varying excitation source. Therefore, the information present in a speech such as message, emotion, language, speaker is due to the combined ef…
Emotion RecognitionLost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largel…
Speech Emotion RecognitionQuestion Answeringvoice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervised speech models. Prior work on singing …
Speech RecognitionA Survey on Paralinguistics in Tamil Speech Processing
Speech carries not only the semantic content but also the paralinguistic information which captures the speaking style. Speaker traits and emotional states affect how words are being spoken. The research on paralinguisti…
Emotion RecognitionSpeaker Identificationspeech-recognitionSpeech Recognition+1Speech Emotion Recognition Considering Local Dynamic Features
Recently, increasing attention has been directed to the study of the speech emotion recognition, in which global acoustic features of an utterance are mostly used to eliminate the content differences. However, the expres…
Emotion RecognitionSpeech Emotion Recognition