paper-with-me

Papers

Learning Disentangled Speech Representations

2023-11-04 · Yusuf Brima, Ulf Krumnack, Simone Pika, Gunther Heidemann

Disentangled representation learning in speech processing has lagged behind other domains, largely due to the lack of datasets with annotated generative factors for robust evaluation. To address this, we propose SynSpeech, a novel large-scale synthetic speech dataset specifically designed to enable research on disentangled speech representations. SynSpeech includes controlled variations in speaker identity, spoken text, and speaking style, with three dataset versions to support experimentation at different levels of complexity. In this study, we present a comprehensive framework to evaluate disentangled representation learning techniques, applying both linear probing and established supervised disentanglement metrics to assess the modularity, compactness, and informativeness of the representations learned by a state-of-the-art model. Using the RAVE model as a test case, we find that SynSpeech facilitates benchmarking across a range of factors, achieving promising disentanglement of simpler features like gender and speaking style, while highlighting challenges in isolating complex attributes like speaker identity. This benchmark dataset and evaluation framework fills a critical gap, supporting the development of more robust and interpretable speech representation learning methods.

📄 PDF Abstract BibTeX arXiv:2311.03389

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDisentanglementInformativenessRepresentation LearningSpeech Representation Learning

Similar Papers 제목 키워드 기반

Towards Learning Fine-Grained Disentangled Representations from Speech

2018-08-08 · Yuan Gong, Christian Poellabauer

Learning disentangled representations of high-dimensional data is currently an active research area. However, compared to the field of computer vision, less work has been done for speech processing. In this paper, we pro…

Representation LearningSpeech Representation Learning

Adversarially learning disentangled speech representations for robust multi-factor voice conversion

2021-01-30 · Jie Wang, Jingbei Li, Xintao Zhao, Zhiyong Wu 외

Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech a…

Representation LearningRhythmSpeech Representation LearningStyle Transfer+1

Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation

2024-11-26 · Pu Wang, Hugo Van hamme

End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this stud…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Investigating Speaker Embedding Disentanglement on Natural Read Speech

2023-08-08 · Michael Kuhlmann, Adrian Meise, Fritz Seebauer, Petra Wagner 외

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…

DisentanglementFairnessRepresentation Learning

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

2021-04-01 · Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov 외

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic inform…

DisentanglementRepresentation LearningResynthesisSpeaker Identification+1