paper-with-me

Papers

Using joint training speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion

2023-07-01 · Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro

Voice conversion systems have made significant advancements in terms of naturalness and similarity in common voice conversion tasks. However, their performance in more complex tasks such as cross-lingual voice conversion and expressive voice conversion remains imperfect. In this study, we propose a novel approach that combines a jointly trained speaker encoder and content features extracted from the cross-lingual speech recognition model Whisper to achieve high-quality cross-lingual voice conversion. Additionally, we introduce a speaker consistency loss to the joint encoder, which improves the similarity between the converted speech and the reference speech. To further explore the capabilities of the joint speaker encoder, we use the phonetic posteriorgram as the content feature, which enables the model to effectively reproduce both the speaker characteristics and the emotional aspects of the reference speech.

📄 PDF Abstract BibTeX arXiv:2307.00393

Code (1)

ConsistencyVC/ConsistencyVC-voive-conversion 공식 구현 pytorch

Tasks

speech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

Building Bilingual and Code-Switched Voice Conversion with Limited Training Data Using Embedding Consistency Loss

2021-04-22 · Yaogen Yang, Haozhe Zhang, Xiaoyi Qin, Shanshan Liang 외

Building cross-lingual voice conversion (VC) systems for multiple speakers and multiple languages has been a challenging task for a long time. This paper describes a parallel non-autoregressive network to achieve bilingu…

Voice CloningVoice Conversion

DNCASR: End-to-End Training for Speaker-Attributed ASR

2025-06-02 · Xianrui Zheng, Chao Zhang, Philip C. Woodland

This paper introduces DNCASR, a novel end-to-end trainable system designed for joint neural speaker clustering and automatic speech recognition (ASR), enabling speaker-attributed transcription of long multi-party meeting…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement

2024-09-02 · Tathagata Bandyopadhyay

Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a trans…

Target Speaker Extraction

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

2025-03-13 · Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…

Speaker Identificationspeech-recognitionSpeech Recognition

Cycle-consistency training for end-to-end speech recognition

2018-11-02 · Takaaki Hori, Ramon Astudillo, Tomoki Hayashi, Yu Zhang 외

This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionar…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+4