paper-with-me

Papers

Adversarially learning disentangled speech representations for robust multi-factor voice conversion

2021-01-30 · Jie Wang, Jingbei Li, Xintao Zhao, Zhiyong Wu, Shiyin Kang, Helen Meng

Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech as speaker and content, lacking controllability on other prosody-related factors. State-of-the-art speech representation learning methods for more speechfactors are using primary disentangle algorithms such as random resampling and ad-hoc bottleneck layer size adjustment,which however is hard to ensure robust speech representationdisentanglement. To increase the robustness of highly controllable style transfer on multiple factors in VC, we propose a disentangled speech representation learning framework based on adversarial learning. Four speech representations characterizing content, timbre, rhythm and pitch are extracted, and further disentangled by an adversarial Mask-And-Predict (MAP)network inspired by BERT. The adversarial network is used tominimize the correlations between the speech representations,by randomly masking and predicting one of the representationsfrom the others. Experimental results show that the proposedframework significantly improves the robustness of VC on multiple factors by increasing the speech quality MOS from 2.79 to3.30 and decreasing the MCD from 3.89 to 3.58.

📄 PDF Abstract BibTeX arXiv:2102.00184

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningRhythmSpeech Representation LearningStyle TransferVoice Conversion

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
WordPiece 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Audio-visual Speech Separation with Adversarially Disentangled Visual Representation

2020-11-29 · Peng Zhang, Jiaming Xu, Jing Shi, Yunzhe Hao 외

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefin…

Speech Separation

Investigating Speaker Embedding Disentanglement on Natural Read Speech

2023-08-08 · Michael Kuhlmann, Adrian Meise, Fritz Seebauer, Petra Wagner 외

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…

DisentanglementFairnessRepresentation Learning

Learning Disentangled Speech Representations

2023-11-04 · Yusuf Brima, Ulf Krumnack, Simone Pika, Gunther Heidemann

Disentangled representation learning in speech processing has lagged behind other domains, largely due to the lack of datasets with annotated generative factors for robust evaluation. To address this, we propose SynSpeec…

BenchmarkingDisentanglementInformativenessRepresentation Learning+1

Powerful Speaker Embedding Training Framework by Adversarially Disentangled Identity Representation

2019-11-27 · Jianwei Tai, Hang Zhou, Qingjia Huang, Xiaoqi Jia

The main challenge of speaker verification in the wild is the interference caused by irrelevant information in speech and the lack of speaker labels in speech datasets. In order to solve the above problems, we propose a …

Speaker Verification

FAVAE: Sequence Disentanglement using Information Bottleneck Principle

2019-02-22 · Masanori Yamada, Heecheol Kim, Kosuke Miyoshi, Hiroshi Yamakawa

We propose the factorized action variational autoencoder (FAVAE), a state-of-the-art generative model for learning disentangled and interpretable representations from sequential data via the information bottleneck withou…

DisentanglementRepresentation Learning