paper-with-me

홈 › Papers

Iteratively Improving Speech Recognition and Voice Conversion

2023-05-24 · Mayank Kumar Singh, Naoya Takahashi, Onoe Naoyuki

Many existing works on voice conversion (VC) tasks use automatic speech recognition (ASR) models for ensuring linguistic consistency between source and converted samples. However, for the low-data resource domains, training a high-quality ASR remains to be a challenging task. In this work, we propose a novel iterative way of improving both the ASR and VC models. We first train an ASR model which is used to ensure content preservation while training a VC model. In the next iteration, the VC model is used as a data augmentation method to further fine-tune the ASR model and generalize it to diverse speakers. By iteratively leveraging the improved ASR model to train VC model and vice-versa, we experimentally show improvement in both the models. Our proposed framework outperforms the ASR and one-shot VC baseline models on English singing and Hindi speech domains in subjective and objective evaluations in low-data resource settings.

📄 PDF Abstract BibTeX arXiv:2305.15055

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

SelfVC: Voice Conversion With Iterative Refinement using Self Transformations

2023-10-14 · Paarth Neekhara, Shehzeen Hussain, Rafael Valle, Boris Ginsburg 외

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled represe…

Self-Supervised LearningSpeaker VerificationSpeech SynthesisVoice Conversion

DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion

2024-06-12 · Ziqian Ning, Shuai Wang, Pengcheng Zhu, Zhichao Wang 외

Streaming voice conversion has become increasingly popular for its potential in real-time applications. The recently proposed DualVC 2 has achieved robust and high-quality streaming voice conversion with a latency of abo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage Modeling+4

Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

2020-09-03 · Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen, Ya-Jun Hu 외

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR and TTS system. In this paper, we present…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+5

Using joint training speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion

2023-07-01 · Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro

Voice conversion systems have made significant advancements in terms of naturalness and similarity in common voice conversion tasks. However, their performance in more complex tasks such as cross-lingual voice conversion…

speech-recognitionSpeech RecognitionVoice Conversion

HiFi-VC: High Quality ASR-Based Voice Conversion

2022-03-31 · A. Kashkin, I. Karpukhin, S. Shishkin

The goal of voice conversion (VC) is to convert input voice to match the target speaker's voice while keeping text and prosody intact. VC is usually used in entertainment and speaking-aid systems, as well as applied for …

speech-recognitionSpeech RecognitionVocal Bursts Intensity PredictionVoice Conversion