paper-with-me

홈 › Papers

Dynamic Layer Customization for Noise Robust Speech Emotion Recognition in Heterogeneous Condition Training

2020-10-21 · Alex Wilf, Emily Mower Provost

Robustness to environmental noise is important to creating automatic speech emotion recognition systems that are deployable in the real world. Prior work on noise robustness has assumed that systems would not make use of sample-by-sample training noise conditions, or that they would have access to unlabelled testing data to generalize across noise conditions. We avoid these assumptions and introduce the resulting task as heterogeneous condition training. We show that with full knowledge of the test noise conditions, we can improve performance by dynamically routing samples to specialized feature encoders for each noise condition, and with partial knowledge, we can use known noise conditions and domain adaptation algorithms to train systems that generalize well to unseen noise conditions. We then extend these improvements to the multimodal setting by dynamically routing samples to maintain temporal ordering, resulting in significant improvements over approaches that do not specialize or generalize based on noise type.

📄 PDF Abstract BibTeX arXiv:2010.11226

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationEmotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

On the Effectiveness of ASR Representations in Real-world Noisy Speech Emotion Recognition

2023-11-13 · Xiaohan Shi, Jiajun He, Xingfeng Li, Tomoki Toda

This paper proposes an efficient attempt to noisy speech emotion recognition (NSER). Conventional NSER approaches have proven effective in mitigating the impact of artificial noise sources, such as white Gaussian noise, …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSelf-Supervised Learning+3

Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing

2026-01-26 · Marouane El Hizabri, Abdelfattah Bezzaz, Ismail Hayoukane, Youssef Taki arxiv

Speech Emotion Recognition systems often use static features like Mel-Frequency Cepstral Coefficients (MFCCs), Zero Crossing Rate (ZCR), and Root Mean Square Energy (RMSE). Because of this, they can misclassify emotions …

Speech Emotion RecognitionEmotion Classification

Investigations on Audiovisual Emotion Recognition in Noisy Conditions

2021-03-02 · Michael Neumann, Ngoc Thang Vu

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition per…

Emotion RecognitionSpeech Emotion Recognition

Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues

2024-02-26 · Tassadaq Hussain, Kia Dashtipour, Yu Tsao, Amir Hussain

In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often f…

DecoderSpeech Enhancement

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

2025-09-20 · Sirui Wang, Andong Chen, Tiejun Zhao arxiv

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural…

Speech Synthesis