paper-with-me

홈 › Papers

Autoregressive Speech Enhancement via Acoustic Tokens

2025-07-17 · Luca Della Libera, Cem Subakan, Mirco Ravanelli

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising alternative for a smooth integration with other modalities. However, research on speech enhancement using discrete representations is still limited. Previous work has mainly focused on semantic tokens, which tend to discard key acoustic details such as speaker identity. Additionally, these studies typically employ non-autoregressive models, assuming conditional independence of outputs and overlooking the potential improvements offered by autoregressive modeling. To address these gaps we: 1) conduct a comprehensive study of the performance of acoustic tokens for speech enhancement, including the effect of bitrate and noise strength; 2) introduce a novel transducer-based autoregressive architecture specifically designed for this task. Experiments on VoiceBank and Libri1Mix datasets show that acoustic tokens outperform semantic tokens in terms of preserving speaker identity, and that our autoregressive approach can further improve performance. Nevertheless, we observe that discrete representations still fall short compared to continuous ones, highlighting the need for further research in this area.

📄 PDF Abstract BibTeX arXiv:2507.12825

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

High-Fidelity Speech Enhancement via Discrete Audio Tokens

2025-10-02 · Luca A. Lanzendörfer, Frédéric Berdoz, Antonis Asonitis, Roger Wattenhofer arxiv

Recent autoregressive transformer-based speech enhancement (SE) methods have shown promising results by leveraging advanced semantic understanding and contextual modeling of speech. However, these approaches often rely o…

Speech Enhancement

A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis

2024-06-18 · Guoqiang Hu, Huaning Tan, Ruilai Li

Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained …

DecoderSpeech Synthesis

MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

2024-09-01 · Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng 외

The recent large-scale text-to-speech (TTS) systems are usually grouped as autoregressive and non-autoregressive systems. The autoregressive systems implicitly model duration but exhibit certain deficiencies in robustnes…

Self-Supervised Learningtext-to-speechText to Speech

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

2025-02-05 · Jixun Yao, Hexin Liu, Chen Chen, Yuchen Hu 외

Semantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns…

Language ModelingLanguage ModellingSpeech Enhancement

UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition

2025-09-18 · Ying Fang, Xiaofei Li arxiv

This paper proposes a unimodal aggregation (UMA) based nonautoregressive model for both English and Mandarin speech recognition. The original UMA explicitly segments and aggregates acoustic frames (with unimodal weights …

Speech Recognition