paper-with-me

Papers

Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks

2024-07-26 · Mahmoud Salhab, Haidar Harmanani

Speech bandwidth expansion is crucial for expanding the frequency range of low-bandwidth speech signals, thereby improving audio quality, clarity and perceptibility in digital applications. Its applications span telephony, compression, text-to-speech synthesis, and speech recognition. This paper presents a novel approach using a high-fidelity generative adversarial network, unlike cascaded systems, our system is trained end-to-end on paired narrowband and wideband speech signals. Our method integrates various bandwidth upsampling ratios into a single unified model specifically designed for speech bandwidth expansion applications. Our approach exhibits robust performance across various bandwidth expansion factors, including those not encountered during training, demonstrating zero-shot capability. To the best of our knowledge, this is the first work to showcase this capability. The experimental results demonstrate that our method outperforms previous end-to-end approaches, as well as interpolation and traditional techniques, showcasing its effectiveness in practical speech enhancement applications.

📄 PDF Abstract BibTeX arXiv:2407.18571

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkSpeech Enhancementspeech-recognitionSpeech RecognitionSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Joint domain adaptation and speech bandwidth extension using time-domain GANs for speaker verification

2022-03-30 · Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez, Najim Dehak

Speech systems developed for a particular choice of acoustic domain and sampling frequency do not translate easily to others. The usual practice is to learn domain adaptation and bandwidth extension models independently.…

Bandwidth ExtensionDomain AdaptationSpeaker Verification

DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers

2025-04-13 · Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar, Tiago H. Falk 외

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approa…

HallucinationSpeech Enhancement

VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

2022-04-12 · Haohe Liu, Xubo Liu, Qiuqiang Kong, Qiao Tian 외

Speech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, speech signals can be degraded by several …

Speech DenoisingSpeech EnhancementVocal Bursts Intensity Prediction

CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space

2026-03-02 · Bowen Zhang, Junchuan Zhao, Ian McLoughlin, Ye Wang 외 arxiv

Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on spectrogram or waveform modeling, which c…

Bandwidth Extension

High-Fidelity Audio Compression with Improved RVQGAN

2023-06-11 · NeurIPS 2023 11 · Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 외

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natur…

Audio CompressionAudio GenerationQuantization