End-to-End Optimized Speech Coding with Deep Neural Networks
Modern compression algorithms are often the result of laborious domain-specific research; industry standards such as MP3, JPEG, and AMR-WB took years to develop and were largely hand-designed. We present a deep neural network model which optimizes all the steps of a wideband speech coding pipeline (compression, quantization, entropy coding, and decompression) end-to-end directly from raw speech data -- no manual feature engineering necessary, and it trains in hours. In testing, our DNN-based coder performs on par with the AMR-WB standard at a variety of bitrates (~9kbps up to ~24kbps). It also runs in realtime on a 3.8GhZ Intel CPU.
Code (0)
등록된 구현이 없습니다.
Tasks
CPUFeature EngineeringQuantizationSimilar Papers 제목 키워드 기반
Lightweight Diffusion-based Framework for Online Imagined Speech Decoding in Aphasia
Individuals with aphasia experience severe difficulty in real-time verbal communication, while most imagined speech decoding approaches remain limited to offline analysis or computationally demanding models. To address t…
Dimensionality ReductionDRED: Deep REDundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder
Despite recent advancements in packet loss concealment (PLC) using deep learning techniques, packet loss remains a significant challenge in real-time speech communication. Redundancy has been used in the past to recover …
Packet Loss ConcealmentEnd-to-End Neural Speech Coding for Real-Time Communications
Deep-learning based methods have shown their advantages in audio coding over traditional ones but limited attention has been paid on real-time communications (RTC). This paper proposes the TFNet, an end-to-end neural spe…
DecoderPacket Loss ConcealmentSpeech EnhancementAn efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks
Auditory front-end is an integral part of a spiking neural network (SNN) when performing auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reco…
Benchmarkingspeech-recognitionSpeech RecognitionUnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units
Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass dir…
DecoderDenoisingSpeech-to-Speech TranslationTranslation+2