VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency
We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus on learning the alignment between the speaker's lip movements and the sounds they generate, we propose to leverage the speaker's face appearance as an additional prior to isolate the corresponding vocal qualities they are likely to produce. Our approach jointly learns audio-visual speech separation and cross-modal speaker embeddings from unlabeled video. It yields state-of-the-art results on five benchmark datasets for audio-visual speech separation and enhancement, and generalizes well to challenging real-world videos of diverse scenarios. Our video results and code: http://vision.cs.utexas.edu/projects/VisualVoice/.
Code (1)
Tasks
Speech SeparationSimilar Papers 제목 키워드 기반
Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss
We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual spee…
Speech SeparationTDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Desig…
speech-recognitionSpeech RecognitionSpeech SeparationAV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an \gls{av} system for speech enhancement, target speaker extraction, and multi-talker speaker se…
Speaker SeparationSpeech EnhancementSpeech SeparationTarget Speaker ExtractionAudio-visual speech separation based on joint feature representation with cross-modal attention
Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward f…
Optical Flow EstimationSpeech SeparationCSLNSpeech: solving extended speech separation problem with the help of Chinese sign language
Previous audio-visual speech separation methods use the synchronization of the speaker's facial movement and speech in the video to supervise the speech separation in a self-supervised way. In this paper, we propose a mo…
Self-Supervised LearningSpeech Separation