paper-with-me

Papers

Disentangled Feature Learning for Real-Time Neural Speech Coding

2022-11-22 · Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned, vector-quantized and coded. In this paper, instead of blind end-to-end learning, we propose to learn disentangled features for real-time neural speech coding. Specifically, more global-like speaker identity and local content features are learned with disentanglement to represent speech. Such a compact feature decomposition not only achieves better coding efficiency by exploiting bit allocation among different features but also provides the flexibility to do audio editing in embedding space, such as voice conversion in real-time communications. Both subjective and objective results demonstrate its coding efficiency and we find that the learned disentangled features show comparable performance on any-to-any voice conversion with modern self-supervised speech representation learning models with far less parameters and low latency, showing the potential of our neural coding framework.

📄 PDF Abstract BibTeX arXiv:2211.11960

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementRepresentation LearningSpeech Representation LearningVoice Conversion

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

Disentangled Speech Representation Learning Based on Factorized Hierarchical Variational Autoencoder with Self-Supervised Objective

2022-04-05 · Yuying Xie, Thomas Arildsen, Zheng-Hua Tan

Disentangled representation learning aims to extract explanatory features or factors and retain salient information. Factorized hierarchical variational autoencoder (FHVAE) presents a way to disentangle a speech signal i…

DisentanglementRepresentation LearningSpeaker Recognitionspeech-recognition+3

Coding Speech through Vocal Tract Kinematics

2024-06-18 · Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal 외

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spo…

Voice Conversion

Investigating Speaker Embedding Disentanglement on Natural Read Speech

2023-08-08 · Michael Kuhlmann, Adrian Meise, Fritz Seebauer, Petra Wagner 외

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…

DisentanglementFairnessRepresentation Learning

LFIC-DRASC: Deep Light Field Image Compression Using Disentangled Representation and Asymmetrical Strip Convolution

2024-09-18 · Shiyu Feng, Yun Zhang, Linwei Zhu, Sam Kwong

Light-Field (LF) image is emerging 4D data of light rays that is capable of realistically presenting spatial and angular information of 3D scene. However, the large data volume of LF images becomes the most challenging i…

Image Compression

DSNet: Disentangled Siamese Network with Neutral Calibration for Speech Emotion Recognition

2023-12-25 · Chengxin Chen, Pengyuan Zhang

One persistent challenge in deep learning based speech emotion recognition (SER) is the unconscious encoding of emotion-irrelevant factors (e.g., speaker or phonetic variability), which limits the generalization of SER i…

DisentanglementEmotion RecognitionSpeech Emotion Recognition