paper-with-me

Papers

Efficient Neural and Numerical Methods for High-Quality Online Speech Spectrogram Inversion via Gradient Theorem

2025-05-30 · Andres Fernandez, Juan Azcarreta, Cagdas Bilen, Jesus Monge Alvarez

Recent work in online speech spectrogram inversion effectively combines Deep Learning with the Gradient Theorem to predict phase derivatives directly from magnitudes. Then, phases are estimated from their derivatives via least squares, resulting in a high quality reconstruction. In this work, we introduce three innovations that drastically reduce computational cost, while maintaining high quality: Firstly, we introduce a novel neural network architecture with just 8k parameters, 30 times smaller than previous state of the art. Secondly, increasing latency by 1 hop size allows us to further halve the cost of the neural inference step. Thirdly, we we observe that the least squares problem features a tridiagonal matrix and propose a linear-complexity solver for the least squares step that leverages tridiagonality and positive-semidefiniteness, achieving a speedup of several orders of magnitude. We release samples online.

📄 PDF Abstract BibTeX arXiv:2505.24498

Code (0)

등록된 구현이 없습니다.

Tasks

8k

Similar Papers 제목 키워드 기반

Semi-Blind Source Separation for Nonlinear Acoustic Echo Cancellation

2020-10-25 · Guoliang Cheng, Lele Liao, Hongsheng Chen, Jing Lu

The mismatch between the numerical and actual nonlinear models is a challenge to nonlinear acoustic echo cancellation (NAEC) when the nonlinear adaptive filter is utilized. To alleviate this problem, we combine a basis-g…

Acoustic echo cancellationblind source separation

EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation

2024-06-10 · Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker 외

We release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totaling in 100 hours of clean, anechoic speech data. The dataset co…

Speech Enhancement

Exploring Speech Enhancement for Low-resource Speech Synthesis

2023-09-19 · Zhaoheng Ni, Sravya Popuri, Ning Dong, Kohei Saijo 외

High-quality and intelligible speech is essential to text-to-speech (TTS) model training, however, obtaining high-quality data for low-resource languages is challenging and expensive. Applying speech enhancement on Autom…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+4

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

2025-03-26 · Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 외

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedba…

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

2026-07-22 · Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen 외 hf

Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are r…