paper-with-me

홈 › Papers

Speech Enhancement with Multi-granularity Vector Quantization

2023-02-16 · Xiao-Ying Zhao, Qiu-Shi Zhu, Jie Zhang

With advances in deep learning, neural network based speech enhancement (SE) has developed rapidly in the last decade. Meanwhile, the self-supervised pre-trained model and vector quantization (VQ) have achieved excellent performance on many speech-related tasks, while they are less explored on SE. As it was shown in our previous work that utilizing a VQ module to discretize noisy speech representations is beneficial for speech denoising, in this work we therefore study the impact of using VQ at different layers with different number of codebooks. Different VQ modules indeed enable to extract multiple-granularity speech features. Following an attention mechanism, the contextual features extracted by a pre-trained model are fused with the local features extracted by the encoder, such that both global and local information are preserved to reconstruct the enhanced speech. Experimental results on the Valentini dataset show that the proposed model can improve the SE performance, where the impact of choosing pre-trained models is also revealed.

📄 PDF Abstract BibTeX arXiv:2302.08342

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingQuantizationSpeech DenoisingSpeech Enhancement

Similar Papers 제목 키워드 기반

Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech

2024-02-26 · Szu-Wei Fu, Kuo-Hsuan Hung, Yu Tsao, Yu-Chiang Frank Wang

Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expe…

QuantizationSpeech Enhancement

MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization

2025-03-18 · Binjie Liu, Lina Liu, Sanyi Zhang, Songen Gu 외

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and …

Gesture GenerationQuantization

ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs

2024-10-16 · Rui-Chen Zheng, Hui-Peng Du, Xiao-Hang Jiang, Yang Ai 외

Current neural audio codecs typically use residual vector quantization (RVQ) to discretize speech signals. However, they often experience codebook collapse, which reduces the effective codebook size and leads to suboptim…

DiversityOnline ClusteringQuantizationtext-to-speech+1

A study on speech enhancement using exponent-only floating point quantized neural network (EOFP-QNN)

2018-08-17 · Yi-Te Hsu, Yu-Chen Lin, Szu-Wei Fu, Yu Tsao 외

Numerous studies have investigated the effectiveness of neural network quantization on pattern classification tasks. The present study, for the first time, investigated the performance of speech enhancement (a regression…

QuantizationregressionSpeech Enhancement

Speech Enhancement Using Self-Supervised Pre-Trained Model and Vector Quantization

2022-09-28 · Xiao-Ying Zhao, Qiu-Shi Zhu, Jie Zhang

With the development of deep learning, neural network-based speech enhancement (SE) models have shown excellent performance. Meanwhile, it was shown that the development of self-supervised pre-trained models can be appli…

DecoderDenoisingQuantizationSpeech Enhancement