Papers Speech Extraction
“Speech Extraction” 태그가 달린 논문 55편 · 필터 해제
Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance …
Speech ExtractionIsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. W…
Speech ExtractionVorTEX: Various overlap ratio for Target speech EXtraction
Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into be…
Speech ExtractionLightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivati…
Speech EnhancementSpeech ExtractionSpeech SeparationELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and p…
Speech ExtractionNeural Speech Extraction with Human Feedback
We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The ref…
Speech ExtractionTF-MLPNet: Tiny Real-Time Neural Speech Separation
Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelera…
Speech ExtractionSpeech SeparationIncorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semanti…
Speech ExtractionTarget Speaker ExtractionInter-Speaker Relative Cues for Text-Guided Target Speech Extraction
We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relativ…
AttributeSpeech ExtractionSoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent a…
Speech ExtractionSpeech SeparationSingle-Channel Target Speech Extraction Utilizing Distance and Room Clues
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which c…
Speech ExtractionSpeech SeparationSonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce Sonic…
Speech ExtractionContextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike …
Speech ExtractionEnhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spect…
Speech ExtractionBeyond Speaker Identity: Text Guided Target Speech Extraction
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided…
Speech ExtractionSpeech SeparationDistance Based Single-Channel Target Speech Extraction
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological…
Speech ExtractionInvestigation of Speaker Representation for Target-Speaker Speech Processing
Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting infor…
Action DetectionActivity DetectionAutomatic Speech RecognitionSpeaker Recognition+4Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any v…
Bandwidth ExtensionDenoisingSpeech DenoisingSpeech Extraction+1Look Once to Hear: Target Speech Hearing with Noisy Examples
In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achieves this capability, enabling target spe…
CPUSpeech ExtractionSeparate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the chall…
Speech Extraction