paper-with-me

Papers

Speaker-independent neural formant synthesis

2023-06-02 · Pablo Pérez Zarazaga, Zofia Malisz, Gustav Eje Henter, Lauri Juvela

We describe speaker-independent speech synthesis driven by a small set of phonetically meaningful speech parameters such as formant frequencies. The intention is to leverage deep-learning advances to provide a highly realistic signal generator that includes control affordances required for stimulus creation in the speech sciences. Our approach turns input speech parameters into predicted mel-spectrograms, which are rendered into waveforms by a pre-trained neural vocoder. Experiments with WaveNet and HiFi-GAN confirm that the method achieves our goals of accurate control over speech parameters combined with high perceptual audio quality. We also find that the small set of phonetically relevant speech parameters we use is sufficient to allow for speaker-independent synthesis (a.k.a. universal vocoding).

📄 PDF Abstract BibTeX arXiv:2306.01957

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Methods 이 논문이 사용한 방법론

Mixture of Logistic Distributions 설명 없음
Dilated Causal Convolution A Dilated Causal Convolution is a causal convolution where the filter is applied over an area larger than its length by…
WaveNet WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies…
HiFi-GAN HiFi-GAN is a generative adversarial network for speech synthesis. HiFi-GAN consists of one generator and two discriminators: multi-scale and multi-period discriminators. The…

Similar Papers 제목 키워드 기반

Evolution of Voices in French Audiovisual Media Across Genders and Age in a Diachronic Perspective

2024-04-24 · Albert Rilliard, David Doukhan, Rémi Uro, Simon Devauchelle

We present a diachronic acoustic analysis of the voice of 1023 speakers from French media archives. The speakers are spread across 32 categories based on four periods (years 1955/56, 1975/76, 1995/96, 2015/16), four age …

Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation

2024-09-14 · Changjin Han, Seokgi Lee, Gyuhyeon Nam, Gyeongsu Chae

Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary …

Speech Synthesistext-to-speechText to Speech

Improving speaker de-identification with functional data analysis of f0 trajectories

2022-03-31 · Lauri Tavi, Tomi Kinnunen, Rosa González Hautamäki

Due to a constantly increasing amount of speech data that is stored in different types of databases, voice privacy has become a major concern. To respond to such concern, speech researchers have developed various methods…

De-identification

FastPitchFormant: Source-filter based Decomposed Modeling for Speech Synthesis

2021-06-29 · Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim 외

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized spee…

Speech Synthesistext-to-speechText to Speech

Decomposed Temporal Dynamic CNN: Efficient Time-Adaptive Network for Text-Independent Speaker Verification Explained with Speaker Activation Map

2022-03-29 · Seong-Hu Kim, Hyeonuk Nam, Yong-Hwa Park

To extract accurate speaker information for text-independent speaker verification, temporal dynamic CNNs (TDY-CNNs) adapting kernels to each time bin was proposed. However, model size of TDY-CNN is too large and the adap…

Data AugmentationSpeaker VerificationText-Independent Speaker Verification