paper-with-me

홈 › Papers

Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR

2025-10-29 · Shreyas Gopal, Ashutosh Anshul, Haoyang Li, Yue Heng Yeo, Hexin Liu, Eng Siong Chng arxiv

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works that quantize Whisper embeddings for speech-to-unit modeling, we propose disentangling semantic speech content from background noise in the latent space. Our end-to-end model separates clean speech in the form of codebook tokens, while extracting interpretable noise vectors as quantization residue which are supervised via a lightweight classifier. We show that our approach improves alignment between clean/noisy speech and text, producing speech tokens that display a high degree of noiseinvariance, and improves ASR performance. Keeping Whisper frozen, we show an 82% reduction in error rate compared to Whisper, and 35% improvement over baseline methods on the VBDemand test set. Further analyses show that the learned token space generalizes well to both seen and unseen acoustic conditions.

📄 PDF Abstract BibTeX arXiv:2510.25150

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling

2024-09-13 · Sotirios Karapiperis, Nikolaos Ellinas, Alexandra Vioni, Junkwang Oh 외

Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on…

DecoderDisentanglementQuantization

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

2021-04-01 · Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov 외

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic inform…

DisentanglementRepresentation LearningResynthesisSpeaker Identification+1

Investigating Speaker Embedding Disentanglement on Natural Read Speech

2023-08-08 · Michael Kuhlmann, Adrian Meise, Fritz Seebauer, Petra Wagner 외

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…

DisentanglementFairnessRepresentation Learning

Estimating the Completeness of Discrete Speech Units

2024-09-09 · Sung-Lin Yeh, Hao Tang

Representing speech with discrete units has been widely used in speech codec and speech generation. However, there are several unverified claims about self-supervised discrete units, such as disentangling phonetic and sp…

DisentanglementQuantization

Variable-rate hierarchical CPC leads to acoustic unit discovery in speech

2022-06-05 · Santiago Cuervo, Adrian Łańcucki, Ricard Marxer, Paweł Rychlikowski 외

The success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learni…

Acoustic Unit DiscoveryDisentanglementQuantizationSelf-Supervised Learning+2