paper-with-me

Papers

Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison

2025-07-15 · Andrew Valdivia, Yueming Zhang, Hailu Xu, Amir Ghasemkhani, Xin Qin

This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic deviation between the original and cloned utterances indicate potential mispronunciations. Our method leverages recent advances in voice cloning to generate a synthetic version of the user's voice with proper pronunciation, then performs frame-by-frame comparisons to identify problematic segments. Experimental results demonstrate the effectiveness of this approach in pinpointing specific pronunciation errors without requiring predefined phonetic rules or extensive training data for each target language.

📄 PDF Abstract BibTeX arXiv:2507.10985

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Cloning

Similar Papers 제목 키워드 기반

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

2025-02-08 · Wei Deng, Siyi Zhou, Jingchen Shu, Jinchao Wang 외

Recently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.Here, we introd…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+4

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026

2026-06-05 · Seymanur Akti, Alexander Waibel arxiv

Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task is central to speech translation and is the focus of the IWSLT 2026 C…

Reinforcement Learning

E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis

2025-11-10 · Zhisheng Zhang, Derui Wang, Yifan Mi, Zhiyong Wu 외 arxiv

Recent advancements in speech synthesis technology have enriched our daily lives, with high-quality and human-like audio widely adopted across real-world applications. However, malicious exploitation like voice-cloning f…

Speech RecognitionSpeech Synthesis

Expressive Neural Voice Cloning

2021-01-30 · Paarth Neekhara, Shehzeen Hussain, Shlomo Dubnov, Farinaz Koushanfar 외

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, thes…

Speech SynthesisStyle Transfertext-to-speechText to Speech+1

F5-TTS-RO: Extending F5-TTS to Romanian TTS via Lightweight Input Adaptation

2025-12-13 · Radu-Gabriel Chivereanu, Tiberiu Boros arxiv

This work introduces a lightweight input-level adapter for the F5-TTS model that enables Romanian Language support. To preserve the existing capabilities of the model (voice cloning, English and Chinese support), we keep…