paper-with-me

홈 › Papers

EEG-Derived Voice Signature for Attended Speaker Detection

2023-08-28 · Hongxu Zhu, Siqi Cai, Yidi Jiang, Qiquan Zhang, Haizhou Li

\textit{Objective:} Conventional EEG-based auditory attention detection (AAD) is achieved by comparing the time-varying speech stimuli and the elicited EEG signals. However, in order to obtain reliable correlation values, these methods necessitate a long decision window, resulting in a long detection latency. Humans have a remarkable ability to recognize and follow a known speaker, regardless of the spoken content. In this paper, we seek to detect the attended speaker among the pre-enrolled speakers from the elicited EEG signals. In this manner, we avoid relying on the speech stimuli for AAD at run-time. In doing so, we propose a novel EEG-based attended speaker detection (E-ASD) task. \textit{Methods:} We encode a speaker's voice with a fixed dimensional vector, known as speaker embedding, and project it to an audio-derived voice signature, which characterizes the speaker's unique voice regardless of the spoken content. We hypothesize that such a voice signature also exists in the listener's brain that can be decoded from the elicited EEG signals, referred to as EEG-derived voice signature. By comparing the audio-derived voice signature and the EEG-derived voice signature, we are able to effectively detect the attended speaker in the listening brain. \textit{Results:} Experiments show that E-ASD can effectively detect the attended speaker from the 0.5s EEG decision windows, achieving 99.78\% AAD accuracy, 99.94\% AUC, and 0.27\% EER. \textit{Conclusion:} We conclude that it is possible to derive the attended speaker's voice signature from the EEG signals so as to detect the attended speaker in a listening brain. \textit{Significance:} We present the first proof of concept for detecting the attended speaker from the elicited EEG signals in a cocktail party environment. The successful implementation of E-ASD marks a non-trivial, but crucial step towards smart hearing aids.

📄 PDF Abstract BibTeX arXiv:2308.14774

Code (0)

등록된 구현이 없습니다.

Tasks

EEG

Similar Papers 제목 키워드 기반

NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection

2023-12-12 · Zexu Pan, Gordon Wichern, Francois G. Germain, Sameer Khurana 외

Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recor…

EEG

Closing the Gap between Single-User and Multi-User VoiceFilter-Lite

2022-02-24 · Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He 외

VoiceFilter-Lite is a speaker-conditioned voice separation model that plays a crucial role in improving speech recognition and speaker verification by suppressing overlapping speech from non-target speakers. However, one…

Speaker Verificationspeech-recognitionSpeech Recognition

NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

2023-07-26 · Zexu Pan, Marvin Borsdorf, Siqi Cai, Tanja Schultz 외

Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a stro…

EEG

EEG-informed attended speaker extraction from recorded speech mixtures with application in neuro-steered hearing prostheses

2016-02-18 · Simon Van Eyndhoven, Tom Francart, Alexander Bertrand

OBJECTIVE: We aim to extract and denoise the attended speaker in a noisy, two-speaker acoustic scenario, relying on microphone array recordings from a binaural hearing aid, which are complemented with electroencephalogra…

DenoisingEEGElectroencephalogram (EEG)Speech Separation

Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation

2023-10-18 · Tae Jin Park, He Huang, Coleman Hooper, Nithin Koluguri 외

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its capacity to modulate the distribution of s…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1