Empirical Interpretation of the Relationship Between Speech Acoustic Context and Emotion Recognition
Speech emotion recognition (SER) is vital for obtaining emotional intelligence and understanding the contextual meaning of speech. Variations of consonant-vowel (CV) phonemic boundaries can enrich acoustic context with linguistic cues, which impacts SER. In practice, speech emotions are treated as single labels over an acoustic segment for a given time duration. However, phone boundaries within speech are not discrete events, therefore the perceived emotion state should also be distributed over potentially continuous time-windows. This research explores the implication of acoustic context and phone boundaries on local markers for SER using an attention-based approach. The benefits of using a distributed approach to speech emotion understanding are supported by the results of cross-corpora analysis experiments. Experiments where phones and words are mapped to the attention vectors along with the fundamental frequency to observe the overlapping distributions and thereby the relationship between acoustic context and emotion. This work aims to bridge psycholinguistic theory research with computational modelling for SER.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotional IntelligenceEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
Millimeter Wave (mmWave) radar has emerged as a promising modality for speech sensing, offering advantages over traditional microphones. Prior works have demonstrated that radar captures motion signals related to vocal v…
Speech EnhancementFrom Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly eva…
Scene UnderstandingQuestion AnsweringEmpirical Interpretation of Speech Emotion Perception with Attention Based Model for Speech Emotion Recognition
Speech emotion recognition is essential for obtaining emotional intelligence which affects the understanding of context and meaning of speech. Harmonically structured vowel and consonant sounds add indexical and linguist…
Emotional IntelligenceEmotion RecognitionSpeech Emotion RecognitionTowards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on…
Phoneme RecognitionInterpreting intermediate convolutional layers of generative CNNs trained on waveforms
This paper presents a technique to interpret and visualize intermediate layers in generative CNNs trained on raw speech data in an unsupervised manner. We argue that averaging over feature maps after ReLU activation in e…
Time Series Analysis