paper-with-me

Papers

Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing

2024-09-22 · Wenze Ren, Kuo-Hsuan Hung, Rong Chao, YouJin Li, Hsin-Min Wang, Yu Tsao

This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched training and test data. We introduce a post-processing classifier (PPC) to rectify these erroneous outputs, ensuring that the enhanced speech corresponds accurately to the intended speaker. We also adopt a mixup strategy in PPC training to improve its robustness. Experimental results on the AVSE-challenge dataset show that integrating PPC into the AVSE model can significantly improve AVSE performance, and combining PPC with the AVSE model trained with permutation invariant training (PIT) yields the best performance. The proposed method substantially outperforms the baseline model by a large margin. This work highlights the potential for broader applications across various modalities and architectures, providing a promising direction for future research in this field.

📄 PDF Abstract BibTeX arXiv:2409.14554

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Mixup Mixup is a data augmentation technique that generates a weighted combination of random image pairs from the training data. Given two images and their ground truth labels:…

Similar Papers 제목 키워드 기반

On the Role of Visual Cues in Audiovisual Speech Enhancement

2020-04-25 · Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi 외

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target spee…

Self-Supervised LearningSpeech Enhancement

Real-Time System for Audio-Visual Target Speech Enhancement

2025-09-25 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…

Audio-Visual Speech RecognitionSpeech Enhancement

Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders

2019-08-07 · Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin 외

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech sign…

Speech Enhancement

Robust Unsupervised Audio-visual Speech Enhancement Using a Mixture of Variational Autoencoders

2019-11-10 · Mostafa Sadeghi, Xavier Alameda-Pineda

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised…

Speech Enhancement

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

2022-03-31 · CVPR 2022 1 · Karren Yang, Dejan Markovic, Steven Krenn, Vasu Agrawal 외

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet…

Speech Enhancement