paper-with-me

홈 › Papers

Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention

2024-04-29 · Ruijie Tao, Xinyuan Qian, Yidi Jiang, Junjie Li, Jiadong Wang, Haizhou Li

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization. However, this strategy mainly focuses on the existence of target speech, while ignoring the variations of the noise characteristics, i.e., interference speaker and the background noise. That may result in extracting noisy signals from the incorrect sound source in challenging acoustic situations. To this end, we propose a novel selective auditory attention mechanism, which can suppress interference speakers and non-speech signals to avoid incorrect speaker extraction. By estimating and utilizing the undesired noisy signal through this mechanism, we design an AV-TSE framework named Subtraction-and-ExtrAction network (SEANet) to suppress the noisy signals. We conduct abundant experiments by re-implementing three popular AV-TSE methods as the baselines and involving nine metrics for evaluation. The experimental results show that our proposed SEANet achieves state-of-the-art results and performs well for all five datasets. The code can be found in: https://github.com/TaoRuijie/SEANet.git

📄 PDF Abstract BibTeX arXiv:2404.18501

Code (1)

taoruijie/seanet 공식 구현 pytorch

Tasks

Target Speaker Extraction

Similar Papers 제목 키워드 기반

Multimodal Attention Fusion for Target Speaker Extraction

2021-02-02 · Hiroshi Sato, Tsubasa Ochiai, Keisuke Kinoshita, Marc Delcroix 외

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extractio…

Target Speaker Extraction

MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

2024-12-11 · Junjie Li, Ke Zhang, Shuai Wang, Kong Aik Lee 외

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always avail…

Target Speaker Extraction

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

2025-05-27 · Zexu Pan, Shengkui Zhao, Tingting Wang, Kun Zhou 외

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co…

Audio-Visual Speech Enhancement With Selective Off-Screen Speech Extraction

2023-06-10 · Tomoya Yoshinaga, Keitaro Tanaka, Shigeo Morishima

This paper describes an audio-visual speech enhancement (AV-SE) method that estimates from noisy input audio a mixture of the speech of the speaker appearing in an input video (on-screen target speech) and of a selected …

Computational EfficiencySpeech EnhancementSpeech Extraction

ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction

2025-11-09 · Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng 외 arxiv

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and p…

Speech Extraction