LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, they often correlate poorly with perceptual quality and provide limited interpretability for optimization. This work proposes a reinforcement learning-based AVSE framework with a Large Language Model (LLM)-based interpretable reward model. An audio LLM generates natural language descriptions of enhanced speech, which are converted by a sentiment analysis model into a 1-5 rating score serving as the PPO reward for fine-tuning a pretrained AVSE model. Compared with scalar metrics, LLM-generated feedback is semantically rich and explicitly describes improvements in speech quality. Experiments on the 4th COG-MHEAR AVSE Challenge (AVSEC-4) dataset show that the proposed method outperforms a supervised baseline and a DNSMOS-based RL baseline in PESQ, STOI, neural quality metrics, and subjective listening tests.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningSentiment AnalysisSpeech EnhancementSimilar Papers 제목 키워드 기반
On the Role of Visual Cues in Audiovisual Speech Enhancement
We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target spee…
Self-Supervised LearningSpeech EnhancementReal-Time System for Audio-Visual Target Speech Enhancement
We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…
Audio-Visual Speech RecognitionSpeech EnhancementAudio-visual Speech Enhancement Using Conditional Variational Auto-Encoders
Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech sign…
Speech EnhancementAudio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often f…
DecoderSpeech EnhancementRobust Unsupervised Audio-visual Speech Enhancement Using a Mixture of Variational Autoencoders
Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised…
Speech Enhancement