paper-with-me

Papers

Improve few-shot voice cloning using multi-modal learning

2022-03-18 · Haitong Zhang, Yue Lin

Recently, few-shot voice cloning has achieved a significant improvement. However, most models for few-shot voice cloning are single-modal, and multi-modal few-shot voice cloning has been understudied. In this paper, we propose to use multi-modal learning to improve the few-shot voice cloning performance. Inspired by the recent works on unsupervised speech representation, the proposed multi-modal system is built by extending Tacotron2 with an unsupervised speech representation module. We evaluate our proposed system in two few-shot voice cloning scenarios, namely few-shot text-to-speech(TTS) and voice conversion(VC). Experimental results demonstrate that the proposed multi-modal learning can significantly improve the few-shot voice cloning performance over their counterpart single-modal systems.

📄 PDF Abstract BibTeX arXiv:2203.09708

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to SpeechVoice CloningVoice Conversion

Similar Papers 제목 키워드 기반

MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning

2026-01-04 · Chunyu Qiang, Jun Wang, Xiaopeng Wang, Kang Yin 외 arxiv

Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches …

Video Generation

Multi-modal Adversarial Training for Zero-Shot Voice Cloning

2024-08-28 · John Janiczek, Dading Chong, Dongyang Dai, Arlo Faria 외

A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound nat…

Decodertext-to-speechText to SpeechVoice Cloning

Voice Cloning: Comprehensive Survey

2025-05-01 · Hussam Azzuni, Abdulmotaleb El Saddik

Voice Cloning has rapidly advanced in today's digital world, with many researchers and corporations working to improve these algorithms for various applications. This article aims to establish a standardized terminology …

SurveyVoice Cloning

Meta-Voice: Fast few-shot style transfer for expressive voice cloning using meta learning

2021-11-14 · Songxiang Liu, Dan Su, Dong Yu

The task of few-shot style transfer for voice cloning in text-to-speech (TTS) synthesis aims at transferring speaking styles of an arbitrary source speaker to a target speaker's voice using very limited amount of neutral…

DisentanglementMeta-LearningStyle Transfertext-to-speech+2

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

2026-05-07 · Rixi Xu, Qingyu Liu, Haitao Li, Yushen Chen 외 arxiv

In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using th…

Speech Synthesis