SpEx+: A Complete Time Domain Speaker Extraction Network
이 논문의 초록은 아카이브 스냅샷(papers 덤프)에 포함되어 있지 않습니다. 아래 외부 검색으로 원문을 찾아보세요.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech ExtractionResults from the Paper
| Rank | Task | Dataset | Model | Metrics |
|---|---|---|---|---|
| #1 | Speech Extraction | WSJ0-2mix-extr | SpEx+ (tied) | SI-SDR: 18.20 |
Similar Papers 제목 키워드 기반
Conditional Diffusion Model for Target Speaker Extraction
We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuous-time stochastic diffusion process in t…
modelTarget Speaker ExtractionSpEx: Multi-Scale Time Domain Speaker Extraction Network
Speaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct th…
DecoderMulti-Task LearningL-SpEx: Localized Target Speaker Extraction
Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction…
Target Speaker ExtractionNeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is pos…
EEGMulti-stage Speaker Extraction with Utterance and Frame-Level Reference Signals
Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple…