paper-with-me

Papers

Controllable Accent Normalization via Discrete Diffusion

2026-03-15 · Qibing Bai, Yuhan Du, Tom Ko, Shuai Wang, Yannan Wang, Haizhou Li arxiv

Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control.

📄 PDF Abstract BibTeX arXiv:2603.14275

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data

2026-02-22 · Qibing Bai, Shuhao Shi, Shuai Wang, Yukai Ju 외 arxiv

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synth…

Accent conversion using discrete units with parallel data synthesized from controllable accented TTS

2024-09-30 · Tuan Nam Nguyen, Ngoc Quan Pham, Alexander Waibel

The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity wel…

Data AugmentationSpeech Synthesistext-to-speechText to Speech+1

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

2026-07-04 · Qibing Bai, Shuai Wang, Yuhan Du, Bohan Li 외 arxiv

Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The current techniques either require naturally recorded parallel L1-L2 speech for t…

Reinforcement Learning

Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data

2026-03-08 · Thanathai Lertpetchpun, Thanapat Trachu, Jihwan Lee, Tiantian Feng 외 arxiv

Accent is an integral part of society, reflecting multiculturalism and shaping how individuals express identity. The majority of English speakers are non-native (L2) speakers, yet current Text-To-Speech (TTS) systems pri…

Policy Gradient Guidance Enables Test Time Control

2025-10-02 · Jianing Qi, Hao Tang, Zhigang Zhu arxiv

We introduce Policy Gradient Guidance (PGG), a simple extension of classifier-free guidance from diffusion models to classical policy gradient methods. PGG augments the policy gradient with an unconditional branch and in…

Reinforcement LearningContinuous Control