paper-with-me

Papers

Personalized Adaptation with Pre-trained Speech Encoders for Continuous Emotion Recognition

2023-09-05 · Minh Tran, Yufeng Yin, Mohammad Soleymani

There are individual differences in expressive behaviors driven by cultural norms and personality. This between-person variation can result in reduced emotion recognition performance. Therefore, personalization is an important step in improving the generalization and robustness of speech emotion recognition. In this paper, to achieve unsupervised personalized emotion recognition, we first pre-train an encoder with learnable speaker embeddings in a self-supervised manner to learn robust speech representations conditioned on speakers. Second, we propose an unsupervised method to compensate for the label distribution shifts by finding similar speakers and leveraging their label distributions from the training set. Extensive experimental results on the MSP-Podcast corpus indicate that our method consistently outperforms strong personalization baselines and achieves state-of-the-art performance for valence estimation.

📄 PDF Abstract BibTeX arXiv:2309.02418

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion RecognitionValence Estimation

Similar Papers 제목 키워드 기반

SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation

2025-01-20 · Ziling Huang, Haixin Guan, Haoran Wei, Yanhua Long

Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired sp…

Speaker VerificationSpeech Enhancement

Personalized Speech Recognition for Children with Test-Time Adaptation

2024-09-19 · Zhonghao Shi, Harshvardhan Srivastava, Xuan Shi, Shrikanth Narayanan 외

Accurate automatic speech recognition (ASR) for children is crucial for effective real-time child-AI interaction, especially in educational applications. However, off-the-shelf ASR models primarily pre-trained on adult d…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Personalized Automatic Speech Recognition Trained on Small Disordered Speech Datasets

2021-10-09 · Jimmy Tobin, Katrin Tomanek

This study investigates the performance of personalized automatic speech recognition (ASR) for recognizing disordered speech using small amounts of per-speaker adaptation data. We trained personalized models for 195 indi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting

2023-09-22 · NeurIPS 2023 9 · Sungwon Kim ~Sungwon_Kim2, Kevin J. Shih, Rohan Badlani, Joao Felipe Santos 외

While recent large-scale neural codec language models have shown significant improvement in zero-shot TTS by training on thousands of hours of data, they suffer from drawbacks such as a lack of robustness, slow sampling …

DecoderSpeech Synthesis

LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

2025-05-16 · Rao Ma, Tongzhou Chen, Kartik Audhkhasi, Bhuvana Ramabhadran

Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language processing tasks including Automatic Speech Rec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognition+1