paper-with-me

홈 › Papers

SoundStream: An End-to-End Neural Audio Codec

2021-07-07 · Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi

We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStream relies on a model architecture composed by a fully convolutional encoder/decoder network and a residual vector quantizer, which are trained jointly end-to-end. Training leverages recent advances in text-to-speech and speech enhancement, which combine adversarial and reconstruction losses to allow the generation of high-quality audio content from quantized embeddings. By training with structured dropout applied to quantizer layers, a single model can operate across variable bitrates from 3kbps to 18kbps, with a negligible quality loss when compared with models trained at fixed bitrates. In addition, the model is amenable to a low latency implementation, which supports streamable inference and runs in real time on a smartphone CPU. In subjective evaluations using audio at 24kHz sampling rate, SoundStream at 3kbps outperforms Opus at 12kbps and approaches EVS at 9.6kbps. Moreover, we are able to perform joint compression and enhancement either at the encoder or at the decoder side with no additional latency, which we demonstrate through background noise suppression for speech.

📄 PDF Abstract BibTeX arXiv:2107.03312

Code (6)

google/lyra
kaiidams/soundstream-pytorch pytorch
kyutai-labs/moshi pytorch
lucidrains/audiolm-pytorch pytorch
lucidrains/vector-quantize-pytorch pytorch
wesbz/SoundStream pytorch

Tasks

CPUDecoderSpeech Enhancementtext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

SpectroStream: A Versatile Neural Codec for General Audio

2025-08-07 · Yunpeng Li, Kehang Han, Brian McWilliams, Zalan Borsos 외 arxiv

We propose SpectroStream, a full-band multi-channel neural audio codec. Successor to the well-established SoundStream, SpectroStream extends its capability beyond 24 kHz monophonic audio and enables high-quality reconstr…

HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition

2025-06-03 · Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera 외

Compression-based representations (CBRs) from neural audio codecs such as EnCodec capture intricate acoustic features like pitch and timbre, while representation-learning-based representations (RLRs) from pre-trained mod…

Emotion RecognitionRepresentation LearningSpeech Emotion RecognitionSpeech Representation Learning

StreamVC: Real-Time Low-Latency Voice Conversion

2024-01-05 · Yang Yang, Yury Kartynnik, Yunpeng Li, Jiuqiang Tang 외

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces…

Speech SynthesisVoice Conversion

FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec

2023-09-14 · Zhihao Du, Shiliang Zhang, Kai Hu, Siqi Zheng

This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible training recipes and inference scripts fo…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Synthesis+3

Towards audio language modeling - an overview

2024-02-20 · Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai-Wei Chang 외

Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as suitable tokenizers for converting continu…

Language ModelingLanguage Modelling