paper-with-me

Papers

DRED: Deep REDundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder

2022-12-08 · Jean-Marc Valin, Jan Büthe, Ahmed Mustafa, Michael Klingbeil

Despite recent advancements in packet loss concealment (PLC) using deep learning techniques, packet loss remains a significant challenge in real-time speech communication. Redundancy has been used in the past to recover the missing information during losses. However, conventional redundancy techniques are limited in the maximum loss duration they can cover and are often unsuitable for burst packet loss. We propose a new approach based on a rate-distortion-optimized variational autoencoder (RDO-VAE), allowing us to optimize a deep speech compression algorithm for the task of encoding large amounts of redundancy at very low bitrate. The proposed Deep REDundancy (DRED) algorithm can transmit up to 50x redundancy using less than 32 kb/s. Results show that DRED outperforms the existing Opus codec redundancy. We also demonstrate its benefits when operating in the context of WebRTC.

📄 PDF Abstract BibTeX arXiv:2212.04453

Code (0)

등록된 구현이 없습니다.

Tasks

Packet Loss Concealment

Similar Papers 제목 키워드 기반

High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models

2023-09-27 · Chunyu Qiang, Hao Li, Yixin Tian, Yi Zhao 외

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of disc…

AllSpeech Synthesistext-to-speechText to Speech+1

The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

2026-02-17 · Samir Sadok, Laurent Girin, Xavier Alameda-Pineda arxiv

Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space. As a result, they are poorly robust to a global vari…

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

2023-07-28 · Chunyu Qiang, Hao Li, Hao Ni, He Qu 외

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks …

Language ModelingLanguage ModellingmodelSpeech Synthesis+2

Boosting neural video codecs by exploiting hierarchical redundancy

2022-08-08 · Reza Pourreza, Hoang Le, Amir Said, Guillaume Sautiere 외

In video compression, coding efficiency is improved by reusing pixels from previously decoded frames via motion and residual compensation. We define two levels of hierarchical redundancy in video frames: 1) first-order: …

Video Compression

A Digital Predistortion Scheme Exploiting Degrees-of-Freedom for Massive MIMO Systems

2018-01-18

The primary source of nonlinear distortion in wireless transmitters is the power amplifier (PA). Conventional digital predistortion (DPD) schemes use high-order polynomials to accurately approximate and compensate for th…