paper-with-me

Papers

Semi-FedSER: Semi-supervised Learning for Speech Emotion Recognition On Federated Learning using Multiview Pseudo-Labeling

2022-03-15 · Tiantian Feng, Shrikanth Narayanan

Speech Emotion Recognition (SER) application is frequently associated with privacy concerns as it often acquires and transmits speech data at the client-side to remote cloud platforms for further processing. These speech data can reveal not only speech content and affective information but the speaker's identity, demographic traits, and health status. Federated learning (FL) is a distributed machine learning algorithm that coordinates clients to train a model collaboratively without sharing local data. This algorithm shows enormous potential for SER applications as sharing raw speech or speech features from a user's device is vulnerable to privacy attacks. However, a major challenge in FL is limited availability of high-quality labeled data samples. In this work, we propose a semi-supervised federated learning framework, Semi-FedSER, that utilizes both labeled and unlabeled data samples to address the challenge of limited labeled data samples in FL. We show that our Semi-FedSER can generate desired SER performance even when the local label rate l=20 using two SER benchmark datasets: IEMOCAP and MSP-Improv.

📄 PDF Abstract BibTeX arXiv:2203.08810

Code (1)

usc-sail/fed-ser-semi 공식 구현 pytorch

Tasks

Emotion RecognitionFederated LearningSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations

2022-11-11 · Yoori Oh, Juheon Lee, Yoseob Han, Kyogu Lee

Recent text-to-speech models have reached the level of generating natural speech similar to what humans say. But there still have limitations in terms of expressiveness. The existing emotional speech synthesis models hav…

Emotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech

Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning

2023-10-26 · Xinfa Zhu, Yuke Li, Yi Lei, Ning Jiang 외

This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfe…

Contrastive LearningExpressive Speech SynthesisSpeech Synthesis

End-to-End Emotional Speech Synthesis Using Style Tokens and Semi-Supervised Training

2019-06-26

This paper proposes an end-to-end emotional speech synthesis (ESS) method which adopts global style tokens (GSTs) for semi-supervised training. This model is built based on the GST-Tacotron framework. The style tokens ar…

Emotional Speech SynthesisEmotion RecognitionSpeech Synthesis

Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech

2021-10-08 · Pengfei Wu, Junjie Pan, Chenchang Xu, Junhui Zhang 외

In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due to their deduction ability. In response t…

Emotion InterpretationExpressive Speech SynthesisSpeech Synthesistext-to-speech+1

Emo-StarGAN: A Semi-Supervised Any-to-Many Non-Parallel Emotion-Preserving Voice Conversion

2023-09-14 · Suhita Ghosh, Arnab Das, Yamini Sinha, Ingo Siegert 외

Speech anonymisation prevents misuse of spoken data by removing any personal identifier while preserving at least linguistic content. However, emotion preservation is crucial for natural human-computer interaction. The w…

Voice Conversion