paper-with-me

Papers

VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

2022-04-12 · Haohe Liu, Xubo Liu, Qiuqiang Kong, Qiao Tian, Yan Zhao, DeLiang Wang, Chuanzeng Huang, Yuxuan Wang

Speech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, speech signals can be degraded by several different distortions simultaneously in the real world. It is thus important to extend speech restoration models to deal with multiple distortions. In this paper, we introduce VoiceFixer, a unified framework for high-fidelity speech restoration. VoiceFixer restores speech from multiple distortions (e.g., noise, reverberation, and clipping) and can expand degraded speech (e.g., noisy speech) with a low bandwidth to 44.1 kHz full-bandwidth high-fidelity speech. We design VoiceFixer based on (1) an analysis stage that predicts intermediate-level features from the degraded speech, and (2) a synthesis stage that generates waveform using a neural vocoder. Both objective and subjective evaluations show that VoiceFixer is effective on severely degraded speech, such as real-world historical speech recordings. Samples of VoiceFixer are available at https://haoheliu.github.io/voicefixer.

📄 PDF Abstract BibTeX arXiv:2204.05841

Code (1)

haoheliu/voicefixer 공식 구현 pytorch

Tasks

Speech DenoisingSpeech EnhancementVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

2026-05-14 · Xiaoyu Zhan, Xinyu Fu, Chenghao Yang, Xiaohong Zhang 외 arxiv

Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. However, existing methods are either limited to a single modality for audio …

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

2026-05-28 · Bohan Li, Shi Lian, Hankun Wang, Yiwei Guo 외 arxiv

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these…

Speech Synthesis

UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation

2026-03-02 · Hebeizi Li, Zihao Liang, Benyuan Sun, Zihao Yin 외 arxiv

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap i…

Video Generation

UniTTS: Residual Learning of Unified Embedding Space for Speech Style Control

2021-06-21 · Minsu Kang, Sungjae Kim, Injung Kim

We propose a novel high-fidelity expressive speech synthesis model, UniTTS, that learns and controls overlapping style attributes avoiding interference. UniTTS represents multiple style attributes in a single unified emb…

Expressive Speech SynthesisSpeech Synthesis

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

2025-11-23 · Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu 외 arxiv

Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronizat…

Speech Synthesis