paper-with-me

홈 › Papers

Speech Enhancement Using Continuous Embeddings of Neural Audio Codec

2025-02-22 · Haoyang Li, Jia Qi Yip, Tianyu Fan, Eng Siong Chng

Recent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a novel, efficient SE approach by leveraging the pre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based SE methods, which process discrete speech tokens using Language Models (LMs), we perform SE within the continuous embedding space of the pretrained NAC, which is highly compressed along the time dimension for efficient representation. Our lightweight SE model, optimized through an embedding-level loss, delivers results comparable to SE baselines trained on larger datasets, with a significantly lower real-time factor of 0.005. Additionally, our method achieves a low GMAC of 3.94, reducing complexity 18-fold compared to Sepformer in a simulated cloud-based audio transmission environment. This work highlights a new, efficient NAC-based SE solution, particularly suitable for cloud applications where NAC is used to compress audio before transmission. Copyright 20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

📄 PDF Abstract BibTeX arXiv:2502.16240

Code (0)

등록된 구현이 없습니다.

Tasks

QuantizationSpeech Enhancement

Similar Papers 제목 키워드 기반

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

2023-10-07 · Zhihao Du, JiaMing Wang, Qian Chen, Yunfei Chu 외

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLM…

Audio captioningAutomatic Speech RecognitionEmotion RecognitionLanguage Modelling+14

SoundStream: An End-to-End Neural Audio Codec

2021-07-07 · Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund 외

We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStream relies on a model architecture compose…

CPUDecoderSpeech Enhancementtext-to-speech+1

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

2022-03-31 · CVPR 2022 1 · Karren Yang, Dejan Markovic, Steven Krenn, Vasu Agrawal 외

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet…

Speech Enhancement

CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space

2026-03-02 · Bowen Zhang, Junchuan Zhao, Ian McLoughlin, Ye Wang 외 arxiv

Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on spectrogram or waveform modeling, which c…

Bandwidth Extension

Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising

2025-05-20 · Ye-Xin Lu, Hui-Peng Du, Fei Liu, Yang Ai 외

Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains n…

DecoderDenoisingLanguage ModelingLanguage Modelling+4