paper-with-me

Papers

USEV: Universal Speaker Extraction with Visual Cue

2021-09-30 · Zexu Pan, Meng Ge, Haizhou Li

A speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. However, the target-interference speaker overlapping ratios could vary over a wide range from 0% to 100% in natural speech communication, furthermore, the target speaker could be absent in the speech mixture, the speech mixtures in such universal multi-talker scenarios are described as general speech mixtures. The speaker extraction algorithm requires an auxiliary reference, such as a video recording or a pre-recorded speech, to form top-down auditory attention on the target speaker. We advocate that a visual cue, i.e., lip movement, is more informative than an audio cue, i.e., pre-recorded speech, to serve as the auxiliary reference for speaker extraction in disentangling the target speaker from a general speech mixture. In this paper, we propose a universal speaker extraction network with a visual cue, that works for all multi-talker scenarios. In addition, we propose a scenario-aware differentiated loss function for network training, to balance the network performance over different target-interference speaker pairing scenarios. The experimental results show that our proposed method outperforms various competitive baselines for general speech mixtures in terms of signal fidelity.

📄 PDF Abstract BibTeX arXiv:2109.14831

Code (1)

zexupan/usev 공식 구현 pytorch

Similar Papers 제목 키워드 기반

FuseVis: Interpreting neural networks for image fusion using per-pixel saliency visualization

2020-12-06 · Nishant Kumar, Stefan Gumhold

Image fusion helps in merging two or more images to construct a more informative single fused image. Recently, unsupervised learning based convolutional neural networks (CNN) have been utilized for different types of ima…

Autonomous DrivingMulti-Exposure Image Fusion

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

2026-06-16 · Xingyuming Liu, Ruichun Ma, Heyu Guo, Qixiu Li 외 arxiv

Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for robotics rely solely on RGB observations. This limits their ability to perceive…

USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

2024-09-04 · Bang Zeng, Ming Li

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recogniti…

Speaker RecognitionSpeech SeparationTarget Speaker Extraction

Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection

2025-01-07 · Bang Zeng, Ming Li

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction…

Action DetectionActivity DetectionAutomatic Speech RecognitionMulti-Task Learning+5

MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

2024-12-11 · Junjie Li, Ke Zhang, Shuai Wang, Kong Aik Lee 외

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always avail…

Target Speaker Extraction