paper-with-me

Papers

SpEx+: A Complete Time Domain Speaker Extraction Network

2020-05-10

이 논문의 초록은 아카이브 스냅샷(papers 덤프)에 포함되어 있지 않습니다. 아래 외부 검색으로 원문을 찾아보세요.

🔎 Google Scholar BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Extraction

Results from the Paper

RankTaskDatasetModelMetrics
#1 Speech Extraction WSJ0-2mix-extr SpEx+ (tied) SI-SDR: 18.20

Similar Papers 제목 키워드 기반

Conditional Diffusion Model for Target Speaker Extraction

2023-10-07 · Theodor Nguyen, Guangzhi Sun, Xianrui Zheng, Chao Zhang 외

We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuous-time stochastic diffusion process in t…

modelTarget Speaker Extraction

SpEx: Multi-Scale Time Domain Speaker Extraction Network

2020-04-17 · Cheng-Lin Xu, Wei Rao, Eng Siong Chng, Haizhou Li

Speaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct th…

DecoderMulti-Task Learning

L-SpEx: Localized Target Speaker Extraction

2022-02-21 · Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng 외

Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction…

Target Speaker Extraction

NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention

2024-09-04 · Dashanka De Silva, Siqi Cai, Saurav Pahuja, Tanja Schultz 외

In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is pos…

EEG

Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals

2020-11-19 · Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng 외

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple…