paper-with-me

홈 › Papers

VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency

2021-01-08 · CVPR 2021 1 · Ruohan Gao, Kristen Grauman

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus on learning the alignment between the speaker's lip movements and the sounds they generate, we propose to leverage the speaker's face appearance as an additional prior to isolate the corresponding vocal qualities they are likely to produce. Our approach jointly learns audio-visual speech separation and cross-modal speaker embeddings from unlabeled video. It yields state-of-the-art results on five benchmark datasets for audio-visual speech separation and enhancement, and generalizes well to challenging real-world videos of diverse scenarios. Our video results and code: http://vision.cs.utexas.edu/projects/VisualVoice/.

📄 PDF Abstract BibTeX arXiv:2101.03149

Code (1)

facebookresearch/visualvoice pytorch

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

2021-03-02 · Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka 외

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual spee…

Speech Separation

TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion

2024-01-25 · Samuel Pegg, Kai Li, Xiaolin Hu

Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Desig…

speech-recognitionSpeech RecognitionSpeech Separation

AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling

2024-06-17 · Vahid Ahmadi Kalkhorani, Cheng Yu, Anurag Kumar, Ke Tan 외

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an \gls{av} system for speech enhancement, target speaker extraction, and multi-talker speaker se…

Speaker SeparationSpeech EnhancementSpeech SeparationTarget Speaker Extraction

Audio-visual speech separation based on joint feature representation with cross-modal attention

2022-03-05 · Junwen Xiong, Peng Zhang, Lei Xie, Wei Huang 외

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward f…

Optical Flow EstimationSpeech Separation

CSLNSpeech: solving extended speech separation problem with the help of Chinese sign language

2020-07-21 · Jiasong Wu, Xuan Li, Taotao Li, Fanman Meng 외

Previous audio-visual speech separation methods use the synchronization of the speaker's facial movement and speech in the video to supervise the speech separation in a self-supervised way. In this paper, we propose a mo…

Self-Supervised LearningSpeech Separation