paper-with-me

Papers

SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

2023-08-14 · Xiaofei Wang, Manthan Thakker, Zhuo Chen, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Min Tang, Shujie Liu, Jinyu Li, Takuya Yoshioka

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. See https://aka.ms/speechx for demo samples.

📄 PDF Abstract BibTeX arXiv:2308.06873

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMulti-Task LearningSpeech EnhancementTarget Speaker Extractiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

2023-10-07 · Zhihao Du, JiaMing Wang, Qian Chen, Yunfei Chu 외

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLM…

Audio captioningAutomatic Speech RecognitionEmotion RecognitionLanguage Modelling+14

Ultra-Low-Bitrate Speech Coding with Pretrained Transformers

2022-07-05 · Ali Siahkoohi, Michael Chinen, Tom Denton, W. Bastiaan Kleijn 외

Speech coding facilitates the transmission of speech over low-bandwidth networks with minimal distortion. Neural-network based speech codecs have recently demonstrated significant improvements in quality over traditional…

DecoderInductive Bias

TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling

2025-08-22 · Yuancheng Wang, Dekun Chen, Xueyao Zhang, Junan Zhang 외 arxiv

Speech tokenizers serve as foundational components for speech language models, yet current designs exhibit several limitations, including: 1) dependence on multi-layer residual vector quantization structures or high fram…

TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

2025-02-13 · KyungSu Kim, Junghyun Koo, Sungho Lee, Haesun Joung 외

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this a…

Audio GenerationDecoderMusic GenerationText-to-Music Generation+2

TS3-Codec: Transformer-Based Simple Streaming Single Codec

2024-11-27 · Haibin Wu, Naoyuki Kanda, Sefik Emre Eskimez, Jinyu Li

Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstream NAC models are predominantly convoluti…

Audio Compression