paper-with-me

홈 › Papers

Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS

2024-10-19 · Tuan Nam Nguyen, Seymanur Aktı, Ngoc Quan Pham, Alexander Waibel

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence, we developed a new AC approach that not only focuses on accent conversion but also improves pronunciation of non-native accented speaker. By providing the non-native audio and the corresponding transcript, we generate the ideal ground-truth audio with native-like pronunciation with original duration and prosody. This ground-truth data aids the model in learning a direct mapping between accented and native speech. We utilize the end-to-end VITS framework to achieve high-quality waveform reconstruction for the AC task. As a result, our system not only produces audio that closely resembles native accents and while retaining the original speaker's identity but also improve pronunciation, as demonstrated by evaluation results.

📄 PDF Abstract BibTeX arXiv:2410.14997

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement

2025-06-19 · Tuan-Nam Nguyen, Ngoc-Quan Pham, Seymanur Aktı, Alexander Waibel

We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronunciation. Our approach enables stream pr…

text-to-speechText to Speech

Synthetic Cross-accent Data Augmentation for Automatic Speech Recognition

2023-03-01 · Philipp Klumpp, Pooja Chitkara, Leda Sari, Prashant Serai 외

The awareness for biased ASR datasets or models has increased notably in recent years. Even for English, despite a vast amount of available training data, systems perform worse for non-native speakers. In this work, we i…

Automatic Speech RecognitionData Augmentationspeech-recognitionSpeech Recognition

An Investigation of Indian Native Language Phonemic Influences on L2 English Pronunciations

2022-12-19 · Shelly Jain, Priyanshi Pal, Anil Vuppala, Prasanta Ghosh 외

Speech systems are sensitive to accent variations. This is especially challenging in the Indian context, with an abundance of languages but a dearth of linguistic studies characterising pronunciation variations. The grow…

High-Fidelity Neural Phonetic Posteriorgrams

2024-02-27 · Cameron Churchwell, Max Morrison, Bryan Pardo

A phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due to their ability to disentangle pronunc…

Voice Conversion

Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision

2024-08-19 · Zhijun Jia, Huaying Xue, Xiulian Peng, Yan Lu

Low resource of parallel data is the key challenge of accent conversion(AC) problem in which both the pronunciation units and prosody pattern need to be converted. We propose a two-stage generative framework "convert-and…