paper-with-me

홈 › Papers

Combining deep and unsupervised features for multilingual speech emotion recognition

2021-01-10 · International Workshop on pattern recognition for positive teChnology And eldeRly wEllbeing (CARE) 2021 1 · Vincenzo Scotti, Federico Galati, Licia Sbattella, Roberto Tedesco

In this paper we present a Convolutional Neural Network for multilingual emotion recognition from spoken sentences. The purpose of this work was to build a model capable of recognising emotions combining textual and acoustic information compatible with multiple languages. The model we derive has an end-to-end deep architecture, hence it takes raw text and audio data and uses convolutional layers to extract a hierarchy of classification features. Moreover, we show how the trained model achieves good performances in different languages thanks to the usage of multilingual unsupervised textual features. As an additional remark, it is worth to mention that our solution does not require text and audio to be word- or phoneme-aligned. The proposed model, PATHOSnet, was trained and evaluated on multiple corpora with different spoken languages (IEMOCAP, EmoFilm, SES and AESI). Before training, we tuned the hyper-parameters solely on the IEMOCAP corpus, which offers realistic audio recording and transcription of sentences with emotional content in English. The final model turned out to provide state-of-the-art performances on some of the selected data sets on the four considered emotions.

📄 PDF Abstract BibTeX

Code (1)

vincenzo-scotti/workingage_voice_service

Tasks

Emotion RecognitionMultimodal Emotion Recognition

Similar Papers 제목 키워드 기반

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

Multilingual Part-of-Speech Tagging: Two Unsupervised Approaches

2014-01-15 · Tahira Naseem, Benjamin Snyder, Jacob Eisenstein, Regina Barzilay

We demonstrate the effectiveness of multilingual learning for unsupervised part-of-speech tagging. The central assumption of our work is that by combining cues from multiple languages, the structure of each becomes more …

Part-Of-Speech TaggingTAGUnsupervised Part-Of-Speech TaggingVocal Bursts Valence Prediction

Disentangling Prosody Representations with Unsupervised Speech Reconstruction

2022-12-14 · Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin 외

Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DisentanglementEmotion Recognition+6

CLARA: Multilingual Contrastive Learning for Audio Representation Acquisition

2023-10-18 · Kari A Noriy, Xiaosong Yang, Marcin Budka, Jian Jun Zhang

Multilingual speech processing requires understanding emotions, a task made difficult by limited labelled data. CLARA, minimizes reliance on labelled data, enhancing generalization across languages. It excels at fosterin…

Audio ClassificationContrastive LearningCross-Lingual TransferData Augmentation+6

Multilingual and Unsupervised Subword Modeling for Zero-Resource Languages

2018-11-09 · Enno Hermann, Herman Kamper, Sharon Goldwater

Subword modeling for zero-resource languages aims to learn low-level representations of speech audio without using transcriptions or other resources from the target language (such as text corpora or pronunciation diction…

Clustering