paper-with-me

홈 › Papers

WASE: Learning When to Attend for Speaker Extraction in Cocktail Party Environments

2021-06-13 · Yunzhe Hao, Jiaming Xu, Peng Zhang, Bo Xu

In the speaker extraction problem, it is found that additional information from the target speaker contributes to the tracking and extraction of the target speaker, which includes voiceprint, lip movement, facial expression, and spatial information. However, no one cares for the cue of sound onset, which has been emphasized in the auditory scene analysis and psychology. Inspired by it, we explicitly modeled the onset cue and verified the effectiveness in the speaker extraction task. We further extended to the onset/offset cues and got performance improvement. From the perspective of tasks, our onset/offset-based model completes the composite task, a complementary combination of speaker extraction and speaker-dependent voice activity detection. We also combined voiceprint with onset/offset cues. Voiceprint models voice characteristics of the target while onset/offset models the start/end information of the speech. From the perspective of auditory scene analysis, the combination of two perception cues can promote the integrity of the auditory object. The experiment results are also close to state-of-the-art performance, using nearly half of the parameters. We hope that this work will inspire communities of speech processing and psychology, and contribute to communication between them. Our code will be available in https://github.com/aispeech-lab/wase/.

📄 PDF Abstract BibTeX arXiv:2106.07016

Code (1)

aispeech-lab/wase 공식 구현 pytorch

Tasks

Action DetectionActivity Detection

Similar Papers 제목 키워드 기반

NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention

2024-09-04 · Dashanka De Silva, Siqi Cai, Saurav Pahuja, Tanja Schultz 외

In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is pos…

EEG

NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

2023-07-26 · Zexu Pan, Marvin Borsdorf, Siqi Cai, Tanja Schultz 외

Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a stro…

EEG

Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

2023-10-11 · Xiang Hao, Jibin Wu, Jianwei Yu, Chenglin Xu 외

Humans can easily isolate a single speaker from a complex acoustic environment, a capability referred to as the "Cocktail Party Effect." However, replicating this ability has been a significant challenge in the field of …

Language ModellingLarge Language ModelTarget Speaker Extraction

EEG-Derived Voice Signature for Attended Speaker Detection

2023-08-28 · Hongxu Zhu, Siqi Cai, Yidi Jiang, Qiquan Zhang 외

\textit{Objective:} Conventional EEG-based auditory attention detection (AAD) is achieved by comparing the time-varying speech stimuli and the elicited EEG signals. However, in order to obtain reliable correlation values…

EEG

Neural Target Speech Extraction: An Overview

2023-01-31 · Katerina Zmolikova, Marc Delcroix, Tsubasa Ochiai, Keisuke Kinoshita 외

Humans can listen to a target speaker even in challenging acoustic conditions that have noise, reverberation, and interfering speakers. This phenomenon is known as the cocktail-party effect. For decades, researchers have…

Speech Extraction