Speech Extraction
1개 벤치마크 · 논문 55편 · 이 태스크의 논문 보기 →
Benchmarks
WSJ0-2mix-extr
Most implemented
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
Beyond Speaker Identity: Text Guided Target Speech Extraction
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
Look Once to Hear: Target Speech Hearing with Noisy Examples
Neural Target Speech Extraction: An Overview
Papers
Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance …
Speech ExtractionIsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. W…
Speech ExtractionVorTEX: Various overlap ratio for Target speech EXtraction
Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into be…
Speech ExtractionLightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivati…
Speech EnhancementSpeech ExtractionSpeech SeparationELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and p…
Speech ExtractionNeural Speech Extraction with Human Feedback
We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The ref…
Speech Extraction