paper-with-me

홈 › Papers

Online Audio-Visual Autoregressive Speaker Extraction

2025-06-02 · Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang, Kun Zhou, Yukun Ma, Bin Ma

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based on depth-wise separable convolution. Then, we propose a lightweight autoregressive acoustic encoder to serve as the second cue, to actively explore the information in the separated speech signal from past steps. Scenario-wise, for the first time, we study how the algorithm performs when there is a change in focus of attention, i.e., the target speaker. Experimental results on LRS3 datasets show that our visual frontend performs comparably to the previous state-of-the-art on both SkiM and ConvTasNet audio backbones with only 0.1 million network parameters and 2.1 MACs per second of processing. The autoregressive acoustic encoder provides an additional 0.9 dB gain in terms of SI-SNRi, and its momentum is robust against the change in attention.

📄 PDF Abstract BibTeX arXiv:2506.01270

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ConvTasNet 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras

2019-12-05 · Ander Arriandiaga, Giovanni Morrone, Luca Pasa, Leonardo Badino 외

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is th…

Optical Flow EstimationSpeech Separation

Multimodal Attention Fusion for Target Speaker Extraction

2021-02-02 · Hiroshi Sato, Tsubasa Ochiai, Keisuke Kinoshita, Marc Delcroix 외

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extractio…

Target Speaker Extraction

MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

2024-12-11 · Junjie Li, Ke Zhang, Shuai Wang, Kong Aik Lee 외

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always avail…

Target Speaker Extraction

Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors

2023-09-25 · Di Liang, Nian Shao, Xiaofei Li

This work proposes a frame-wise online/streaming end-to-end neural diarization (FS-EEND) method in a frame-in-frame-out fashion. To frame-wisely detect a flexible number of speakers and extract/update their corresponding…

Decoderspeaker-diarizationSpeaker Diarization

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

2025-05-27 · Zexu Pan, Shengkui Zhao, Tingting Wang, Kun Zhou 외

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co…