paper-with-me

Papers

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

2025-11-15 · Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath arxiv

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/.

📄 PDF Abstract BibTeX arXiv:2511.12347

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

2024-03-25 · Puyuan Peng, Po-Yao Huang, Shang-Wen Li, Abdelrahman Mohamed 외

We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts. V…

DecoderLanguage ModelingLanguage Modellingtext-to-speech+1

VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models

2025-04-03 · Kim Sung-Bin, Jeongsoo Choi, Puyuan Peng, Joon Son Chung 외

We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting v…

Speech Synthesis

Zero-Shot Text-to-Speech for Vietnamese

2025-06-02 · Thi Vu, Linh The Nguyen, Dat Quoc Nguyen

This paper introduces PhoAudiobook, a newly curated dataset comprising 941 hours of high-quality audio for Vietnamese text-to-speech. Using PhoAudiobook, we conduct experiments on three leading zero-shot TTS models: VALL…

text-to-speechText to Speech

Your Voice Cloning System is Secretly a Voice Anonymizer

2026-08-27 · Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu arxiv

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speec…

Voice Conversion

Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba

2025-10-06 · Baher Mohammad, Magauiya Zhussip, Stamatios Lefkimmiatis arxiv

We introduce MAVE (Mamba with Cross-Attention for Voice Editing and Synthesis), a novel autoregressive architecture for text-conditioned voice editing and high-fidelity text-to-speech (TTS) synthesis, built on a cross-at…