paper-with-me

Papers

ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding Inpainting

2022-10-31 · Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video clip can serve as the auxiliary reference. The use of visual cue is not only feasible, but also effective due to its noise robustness, and becoming popular. However, it is difficult to guarantee that such parallel visual cue is always available in real-world applications where visual occlusion or intermittent communication can occur. In this paper, we study the audio-visual speaker extraction algorithms with intermittent visual cue. We propose a joint speaker extraction and visual embedding inpainting framework to explore the mutual benefits. To encourage the interaction between the two tasks, they are performed alternately with an interlacing structure and optimized jointly. We also propose two types of visual inpainting losses and study our proposed method with two types of popularly used visual embeddings. The experimental results show that we outperform the baseline in terms of signal quality, perceptual quality, and intelligibility.

📄 PDF Abstract BibTeX arXiv:2211.00109

Code (1)

zexupan/imaginenet 공식 구현 pytorch

Tasks

Target Speaker Extraction

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

ImagineNet: Restyling Apps Using Neural Style Transfer

2020-01-14 · Michael H. Fischer, Richard R. Yang, Monica S. Lam

This paper presents ImagineNet, a tool that uses a novel neural style transfer model to enable end-users and app developers to restyle GUIs using an image of their choice. Former neural style transfer techniques are inad…

Style Transfer

MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

2024-12-11 · Junjie Li, Ke Zhang, Shuai Wang, Kong Aik Lee 외

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always avail…

Target Speaker Extraction

Multimodal Attention Fusion for Target Speaker Extraction

2021-02-02 · Hiroshi Sato, Tsubasa Ochiai, Keisuke Kinoshita, Marc Delcroix 외

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extractio…

Target Speaker Extraction

USEV: Universal Speaker Extraction with Visual Cue

2021-09-30 · Zexu Pan, Meng Ge, Haizhou Li

A speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. H…

L-SpEx: Localized Target Speaker Extraction

2022-02-21 · Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng 외

Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction…

Target Speaker Extraction