paper-with-me

Papers

Iterative autoregression: a novel trick to improve your low-latency speech enhancement model

2022-11-03 · Pavel Andreev, Nicholas Babaev, Azat Saginbaev, Ivan Shchekotov, Aibek Alanov

Streaming models are an essential component of real-time speech enhancement tools. The streaming regime constrains speech enhancement models to use only a tiny context of future information. As a result, the low-latency streaming setup is generally considered a challenging task and has a significant negative impact on the model's quality. However, the sequential nature of streaming generation offers a natural possibility for autoregression, that is, utilizing previous predictions while making current ones. The conventional method for training autoregressive models is teacher forcing, but its primary drawback lies in the training-inference mismatch that can lead to a substantial degradation in quality. In this study, we propose a straightforward yet effective alternative technique for training autoregressive low-latency speech enhancement models. We demonstrate that the proposed approach leads to stable improvement across diverse architectures and training scenarios.

📄 PDF Abstract BibTeX arXiv:2211.01751

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

2026-08-09 · Yuqian Zhang, Yao Shi, Kexin Huang, Botian Jiang 외 arxiv

Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent speaker identity, and low-latency respons…

Speech Synthesis

Boosted p-Values for High-Dimensional Vector Autoregression

2022-11-04 · Xiao Huang

Assessing the statistical significance of parameter estimates is an important step in high-dimensional vector autoregression modeling. Using the least-squares boosting method, we compute the p-value for each selected par…

Time SeriesTime Series AnalysisvalidVocal Bursts Intensity Prediction

Transformer tricks: Precomputing the first layer

2024-02-20 · Nils Graef

This micro-paper describes a trick to speed up inference of transformers with RoPE (such as LLaMA, Mistral, PaLM, and Gemma). For these models, a large portion of the first transformer layer can be precomputed, which res…

Fully Non-autoregressive Neural Machine Translation: Tricks of the Trade

2020-12-31 · Findings (ACL) 2021 8 · Jiatao Gu, Xiang Kong

Fully non-autoregressive neural machine translation (NAT) is proposed to simultaneously predict tokens with single forward of neural networks, which significantly reduces the inference latency at the expense of quality d…

Machine TranslationTranslation

Hybrid Frequency Transmission for Upload Latency Minimization of IoT Devices in HSR Scenario Aided by Intelligent Reflecting Surfaces

2025-02-18 · Tianyou Li, Tonghua Wei, Dapeng Li

The explosively growing demand for Internet of Things (IoT) in high-speed railway (HSR) scenario has attracted a lot of attention amongst researchers. However, limited IoT device (IoTD) batteries and large information up…