paper-with-me

홈 › Papers

Acoustic and linguistic representations for speech continuous emotion recognition in call center conversations

2023-10-06 · Manon Macary, Marie Tahon, Yannick Estève, Daniel Luzzati

The goal of our research is to automatically retrieve the satisfaction and the frustration in real-life call-center conversations. This study focuses an industrial application in which the customer satisfaction is continuously tracked down to improve customer services. To compensate the lack of large annotated emotional databases, we explore the use of pre-trained speech representations as a form of transfer learning towards AlloSat corpus. Moreover, several studies have pointed out that emotion can be detected not only in speech but also in facial trait, in biological response or in textual information. In the context of telephone conversations, we can break down the audio information into acoustic and linguistic by using the speech signal and its transcription. Our experiments confirms the large gain in performance obtained with the use of pre-trained features. Surprisingly, we found that the linguistic content is clearly the major contributor for the prediction of satisfaction and best generalizes to unseen data. Our experiments conclude to the definitive advantage of using CamemBERT representations, however the benefit of the fusion of acoustic and linguistic modalities is not as obvious. With models learnt on individual annotations, we found that fusion approaches are more robust to the subjectivity of the annotation task. This study also tackles the problem of performances variability and intends to estimate this variability from different views: weights initialization, confidence intervals and annotation subjectivity. A deep analysis on the linguistic content investigates interpretable factors able to explain the high contribution of the linguistic modality for this task.

📄 PDF Abstract BibTeX arXiv:2310.04481

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Empirical Interpretation of the Relationship Between Speech Acoustic Context and Emotion Recognition

2023-06-30 · Anna Ollerenshaw, Md Asif Jalal, Rosanna Milner, Thomas Hain

Speech emotion recognition (SER) is vital for obtaining emotional intelligence and understanding the contextual meaning of speech. Variations of consonant-vowel (CV) phonemic boundaries can enrich acoustic context with l…

Emotional IntelligenceEmotion RecognitionSpeech Emotion Recognition

Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration

2026-01-23 · Esther Sun, Abinay Reddy Naini, Carlos Busso arxiv

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization.…

Speech Emotion Recognition

On the Impact of Word Error Rate on Acoustic-Linguistic Speech Emotion Recognition: An Update for the Deep Learning Era

2021-04-20 · Shahin Amiriparian, Artem Sokolov, Ilhan Aslan, Lukas Christ 외

Text encodings from automatic speech recognition (ASR) transcripts and audio representations have shown promise in speech emotion recognition (SER) ever since. Yet, it is challenging to explain the effect of each informa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSelf-Supervised Learning+3

On the use of Self-supervised Pre-trained Acoustic and Linguistic Features for Continuous Speech Emotion Recognition

2020-11-18 · Manon Macary, Marie Tahon, Yannick Estève, Anthony Rousseau

Pre-training for feature extraction is an increasingly studied approach to get better continuous representations of audio and text content. In the present work, we use wav2vec and camemBERT as self-supervised learned mod…

Emotion RecognitionSpeech Emotion Recognition

EMPHASIS: An Emotional Phoneme-based Acoustic Model for Speech Synthesis System

2018-06-26

We present EMPHASIS, an emotional phoneme-based acoustic model for speech synthesis system. EMPHASIS includes a phoneme duration prediction model and an acoustic parameter prediction model. It uses a CBHG-based regressio…

Emotional Speech SynthesisParameter PredictionregressionSpeech Synthesis