paper-with-me

Papers

Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising

2025-05-20 · Ye-Xin Lu, Hui-Peng Du, Fei Liu, Yang Ai, Zhen-Hua Ling

Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS, to achieve noise-robust zero-shot TTS. The proposed codec denoiser consists of an audio codec, a token denoiser, and an embedding refiner. The token denoiser predicts the first two groups of clean acoustic tokens from the noisy ones, which can serve as the acoustic prompt for LauraTTS to synthesize high-quality personalized speech or be converted to clean speech waveforms through the embedding refiner and codec decoder. Experimental results show that our proposed codec denoiser outperforms state-of-the-art speech enhancement (SE) methods, and the proposed noise-robust LauraTTS surpasses the approach using additional SE models.

📄 PDF Abstract BibTeX arXiv:2505.13830

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDenoisingLanguage ModelingLanguage ModellingLarge Language ModelSpeech Enhancementtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

2025-09-11 · Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy 외 arxiv

Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregressive models suffer from high latency, wh…

VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

2024-06-12 · Bing Han, Long Zhou, Shujie Liu, Sanyuan Chen 외

With the help of discrete neural audio codecs, large language models (LLM) have increasingly been recognized as a promising methodology for zero-shot Text-to-Speech (TTS) synthesis. However, sampling based decoding strat…

QuantizationSpeech Synthesistext-to-speechText to Speech+1

Drift-Augmented Scoring: Text-Derived Noise Robustness for Zero-Shot Audio-Language Classification

2026-06-03 · Tu Vo, Sheir Zaheer, Chan Y. Park arxiv

Contrastive audio-language models such as CLAP enable zero-shot audio classification: a sound is labelled by matching its embedding to text prompt embeddings, with no labelled audio. This matching breaks down under acous…

Audio Classification

High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model

2024-06-25 · Joun Yeop Lee, Myeonghun Jeong, Minchan Kim, Ji-Hyun Lee 외

We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting mod…

Computational EfficiencyLanguage ModelingLanguage ModellingSpeech Synthesis+2

R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces

2023-11-15 · Heng-Jui Chang, James Glass

This paper introduces Robust Spin (R-Spin), a data-efficient domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-invariant clust…

ClusteringRepresentation Learning