paper-with-me

Papers

High-Fidelity Speech Enhancement via Discrete Audio Tokens

2025-10-02 · Luca A. Lanzendörfer, Frédéric Berdoz, Antonis Asonitis, Roger Wattenhofer arxiv

Recent autoregressive transformer-based speech enhancement (SE) methods have shown promising results by leveraging advanced semantic understanding and contextual modeling of speech. However, these approaches often rely on complex multi-stage pipelines and low sampling rate codecs, limiting them to narrow and task-specific speech enhancement. In this work, we introduce DAC-SE1, a simplified language model-based SE framework leveraging discrete high-resolution audio representations; DAC-SE1 preserves fine-grained acoustic details while maintaining semantic coherence. Our experiments show that DAC-SE1 surpasses state-of-the-art autoregressive SE methods on both objective perceptual metrics and in a MUSHRA human evaluation. We release our codebase and model checkpoints to support further research in scalable, unified, and high-quality speech enhancement.

📄 PDF Abstract BibTeX arXiv:2510.02187

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers

2025-04-13 · Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar, Tiago H. Falk 외

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approa…

HallucinationSpeech Enhancement

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

2022-03-31 · CVPR 2022 1 · Karren Yang, Dejan Markovic, Steven Krenn, Vasu Agrawal 외

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet…

Speech Enhancement

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

2023-10-07 · Zhihao Du, JiaMing Wang, Qian Chen, Yunfei Chu 외

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLM…

Audio captioningAutomatic Speech RecognitionEmotion RecognitionLanguage Modelling+14

Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders

2025-06-13 · Xingwei Sun, Heinrich Dinkel, Yadong Niu, Linzhang Wang 외

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an e…

Speech Enhancement

From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion

2023-08-02 · NeurIPS 2023 11

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthe…