paper-with-me

Papers

Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP

2024-09-04 · Yisi Liu, Bohan Yu, Drake Lin, Peter Wu, Cheol Jun Cho, Gopala Krishna Anumanchipalli

Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal processing (DDSP) is a parameter-efficient framework for audio synthesis. Therefore, integrating low-dimensional EMA features with DDSP can significantly enhance the computational efficiency of speech synthesis. In this paper, we propose a fast, high-quality, and parameter-efficient DDSP articulatory vocoder that can synthesize speech from EMA, F0, and loudness. We incorporate several techniques to solve the harmonics / noise imbalance problem, and add a multi-resolution adversarial loss for better synthesis quality. Our model achieves a transcription word error rate (WER) of 6.67% and a mean opinion score (MOS) of 3.74, with an improvement of 1.63% and 0.16 compared to the state-of-the-art (SOTA) baseline. Our DDSP vocoder is 4.9x faster than the baseline on CPU during inference, and can generate speech of comparable quality with only 0.4M parameters, in contrast to the 9M parameters required by the SOTA.

📄 PDF Abstract BibTeX arXiv:2409.02451

Code (0)

등록된 구현이 없습니다.

Tasks

Audio SynthesisComputational EfficiencyCPUSpeech Synthesis

Methods 이 논문이 사용한 방법론

DDSP 설명 없음

Similar Papers 제목 키워드 기반

Exploration strategies for articulatory synthesis of complex syllable onsets

2022-04-20 · Daniel R. van Niekerk, Anqi Xu, Branislav Gerazov, Paul K. Krug 외

High-quality articulatory speech synthesis has many potential applications in speech science and technology. However, developing appropriate mappings from linguistic specification to articulatory gestures is difficult an…

Speech Synthesis

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

2026-05-20 · Vinicius Ribeiro, Yves Laprie arxiv

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality ass…

Speech Synthesis

Why can big.bi be changed to bi.gbi? A mathematical model of syllabification and articulatory synthesis

2023-07-05 · Frédéric Berthommier

A simplified model of articulatory synthesis involving four stages is presented. The planning of articulatory gestures is based on syllable graphs with arcs and nodes that are implemented in a complex representation. Thi…

Trajectory Planning

Evaluating Features and Metrics for High-Quality Simulation of Early Vocal Learning of Vowels

2020-05-20 · Branislav Gerazov, Daniel van Niekerk, Anqi Xu, Paul Konstantin Krug 외

The way infants use auditory cues to learn to speak despite the acoustic mismatch of their vocal apparatus is a hot topic of scientific debate. The simulation of early vocal learning using articulatory speech synthesis o…

Speech Synthesis

Coding Speech through Vocal Tract Kinematics

2024-06-18 · Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal 외

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spo…

Voice Conversion