Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
Availability of representations from pre-trained models (PTMs) have facilitated substantial progress in speech emotion recognition (SER). Particularly, representations from PTM trained for paralinguistic speech processing have shown state-of-the-art (SOTA) performance for SER. However, such paralinguistic PTM representations haven't been evaluated for SER in linguistic environments other than English. Also, paralinguistic PTM representations haven't been investigated in benchmarks such as SUPERB, EMO-SUPERB, ML-SUPERB for SER. This makes it difficult to access the efficacy of paralinguistic PTM representations for SER in multiple languages. To fill this gap, we perform a comprehensive comparative study of five SOTA PTM representations. Our results shows that paralinguistic PTM (TRILLsson) representations performs the best and this performance can be attributed to its effectiveness in capturing pitch, tone and other speech characteristics more effectively than other PTM representations.
Code (0)
등록된 구현이 없습니다.
Tasks
AllEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
A Survey on Paralinguistics in Tamil Speech Processing
Speech carries not only the semantic content but also the paralinguistic information which captures the speaking style. Speaker traits and emotional states affect how words are being spoken. The research on paralinguisti…
Emotion RecognitionSpeaker Identificationspeech-recognitionSpeech Recognition+1Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech,…
Emotion RecognitionRhythmSpeaker RecognitionSpeech Emotion Recognition+1On the Emotion Understanding of Synthesized Speech
Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding r…
Speech Emotion RecognitionSpeech SynthesisRecovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization.…
Speech Emotion RecognitionLearning Paralinguistic Features from Audiobooks through Style Voice Conversion
Paralinguistics, the non-lexical components of speech, play a crucial role in human-human interaction. Models designed to recognize paralinguistic information, particularly speech emotion and style, are difficult to trai…
Emotion RecognitionStyle DetectionVoice Conversion