paper-with-me

Papers

Rhythm Modeling for Voice Conversion

2023-07-12 · Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper

Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we introduce Urhythmic-an unsupervised method for rhythm conversion that does not require parallel data or text transcriptions. Using self-supervised representations, we first divide source audio into segments approximating sonorants, obstruents, and silences. Then we model rhythm by estimating speaking rate or the duration distribution of each segment type. Finally, we match the target speaking rate or rhythm by time-stretching the speech segments. Experiments show that Urhythmic outperforms existing unsupervised methods in terms of quality and prosody. Code and checkpoints: https://github.com/bshall/urhythmic. Audio demo page: https://ubisoft-laforge.github.io/speech/urhythmic.

📄 PDF Abstract BibTeX arXiv:2307.06040

Code (1)

bshall/urhythmic 공식 구현 pytorch

Tasks

RhythmVoice Conversion

Similar Papers 제목 키워드 기반

Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech

2025-06-02 · Karl El Hajal, Enno Hermann, Sevada Hovsepyan, Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved A…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythmspeech-recognition+2

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching

2025-06-01 · Jialong Zuo, Shengpeng Ji, Minghui Fang, Mingze Li 외

Zero-Shot Voice Conversion (VC) aims to transform the source speaker's timbre into an arbitrary unseen one while retaining speech content. Most prior work focuses on preserving the source's prosody, while fine-grained ti…

RhythmStyle TransferVoice Conversion

Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity

2024-07-20 · Tianhua Qi, Shiyan Wang, Cheng Lu, Yan Zhao 외

Realistic emotional voice conversion (EVC) aims to enhance emotional diversity of converted audios, making the synthesized voices more authentic and natural. To this end, we propose Emotional Intensity-aware Network (EIN…

DiversityRhythmVoice Conversion

LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling

2024-09-13 · Yubo Huang, Xin Lai, Muyang Ye, Anran Zhu 외

Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer's voice into another while preserving musical elements such as melody, rhythm, and …

CPURhythmVoice Conversion

Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR

2025-01-17 · Karl El Hajal, Enno Hermann, Ajinkya Kulkarni, Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mismatch with typical speech. Unfortunately…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythmspeech-recognition+2