paper-with-me

Papers

GlowVC: Mel-spectrogram space disentangling model for language-independent text-free voice conversion

2022-07-04 · Magdalena Proszewska, Grzegorz Beringer, Daniel Sáez-Trigueros, Thomas Merritt, Abdelhamid Ezzerg, Roberto Barra-Chicote

In this paper, we propose GlowVC: a multilingual multi-speaker flow-based model for language-independent text-free voice conversion. We build on Glow-TTS, which provides an architecture that enables use of linguistic features during training without the necessity of using them for VC inference. We consider two versions of our model: GlowVC-conditional and GlowVC-explicit. GlowVC-conditional models the distribution of mel-spectrograms with speaker-conditioned flow and disentangles the mel-spectrogram space into content- and pitch-relevant dimensions, while GlowVC-explicit models the explicit distribution with unconditioned flow and disentangles said space into content-, pitch- and speaker-relevant dimensions. We evaluate our models in terms of intelligibility, speaker similarity and naturalness for intra- and cross-lingual conversion in seen and unseen languages. GlowVC models greatly outperform AutoVC baseline in terms of intelligibility, while achieving just as high speaker similarity in intra-lingual VC, and slightly worse in the cross-lingual setting. Moreover, we demonstrate that GlowVC-explicit surpasses both GlowVC-conditional and AutoVC in terms of naturalness.

📄 PDF Abstract BibTeX arXiv:2207.01454

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Methods 이 논문이 사용한 방법론

Normalizing Flows Normalizing Flows are a method for constructing complex distributions by transforming a probability density through a series of invertible mappings. By repeatedly applying…
Affine Coupling 설명 없음
Activation Normalization Activation Normalization is a type of normalization used for flow-based generative models; specifically it was introduced in the GLOW
Invertible 1x1 Convolution The Invertible 1x1 Convolution is a type of convolution used in flow-based generative models that reverses the ordering of…
GLOW 설명 없음
Glow-TTS Glow-TTS is a flow-based generative model for parallel TTS that does not require any external aligner. By combining the properties of flows and dynamic programming, the…

Similar Papers 제목 키워드 기반

Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification

2025-06-12 · Peidong Wei, Shiyu Miao, Lin Li

Deep neural networks have been applied to audio spectrograms for respiratory sound classification, but it remains challenging to achieve satisfactory performance due to the scarcity of available data. Moreover, domain mi…

DisentanglementSound Classification

Disentangling Modes and Interference in the Spectrogram of Multicomponent Signals

2025-03-19 · Kévin Polisano, Sylvain Meignen, Nils Laurent, Hubert Leterme

In this paper, we investigate how the spectrogram of multicomponent signals can be decomposed into a mode part and an interference part. We explore two approaches: (i) a variational method inspired by texture-geometry de…

Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations

2024-06-04 · Sarthak Yadav, Zheng-Hua Tan

Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which hav…

Language ModellingMambaState Space Models

Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis

2022-06-23 · Tae-Woo Kim, Min-Su Kang, Gyeong-Hoon Lee

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the …

Generative Adversarial NetworkMulti-Task LearningSinging Voice Synthesis

Three-Dimensional Sparse Random Mode Decomposition for Mode Disentangling with Crossover Instantaneous Frequencies

2025-01-25 · Chen Luo, Tao Chen, Lei Xie, Hongye Su

Sparse random mode decomposition (SRMD) is a novel algorithm that constructs a random time-frequency feature space to sparsely approximate spectrograms, effectively separating modes. However, it fails to distinguish adja…