paper-with-me

홈 › Papers

NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping

2023-09-25 · Jan Büthe, Ahmed Mustafa, Jean-Marc Valin, Karim Helwani, Michael M. Goodwin

Speech codec enhancement methods are designed to remove distortions added by speech codecs. While classical methods are very low in complexity and add zero delay, their effectiveness is rather limited. Compared to that, DNN-based methods deliver higher quality but they are typically high in complexity and/or require delay. The recently proposed Linear Adaptive Coding Enhancer (LACE) addresses this problem by combining DNNs with classical long-term/short-term postfiltering resulting in a causal low-complexity model. A short-coming of the LACE model is, however, that quality quickly saturates when the model size is scaled up. To mitigate this problem, we propose a novel adatpive temporal shaping module that adds high temporal resolution to the LACE model resulting in the Non-Linear Adaptive Coding Enhancer (NoLACE). We adapt NoLACE to enhance the Opus codec and show that NoLACE significantly outperforms both the Opus baseline and an enlarged LACE model at 6, 9 and 12 kb/s. We also show that LACE and NoLACE are well-behaved when used with an ASR system.

📄 PDF Abstract BibTeX arXiv:2309.14521

Code (1)

https://gitlab.xiph.org/xiph/opus

Similar Papers 제목 키워드 기반

Baseline Systems For The 2025 Low-Resource Audio Codec Challenge

2025-09-30 · Yusuf Ziya Isik, Rafał Łaganowski arxiv

The Low-Resource Audio Codec (LRAC) Challenge aims to advance neural audio coding for deployment in resource-constrained environments. The first edition focuses on low-resource neural speech codecs that must operate reli…

Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising

2025-05-20 · Ye-Xin Lu, Hui-Peng Du, Fei Liu, Yang Ai 외

Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains n…

DecoderDenoisingLanguage ModelingLanguage Modelling+4

Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

2026-06-23 · Wangyi Pu, Michele Scarpiniti arxiv

Generative models, particularly diffusion and score-based approaches, have recently achieved strong performance in speech enhancement, but their iterative sampling process limits real-time deployment. Flow Matching offer…

Speech Enhancement

Speech Enhancement Using Continuous Embeddings of Neural Audio Codec

2025-02-22 · Haoyang Li, Jia Qi Yip, Tianyu Fan, Eng Siong Chng

Recent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a novel, efficient SE approach by leveraging …

QuantizationSpeech Enhancement

Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules

2024-10-02 · Hsin-Tien Chiang, Hao Zhang, Yong Xu, Meng Yu 외

In challenging environments with significant noise and reverberation, traditional speech enhancement (SE) methods often lead to over-suppressed speech, creating artifacts during listening and harming downstream tasks per…

QuantizationSpeech Enhancement