paper-with-me

Papers Speech Extraction

“Speech Extraction” 태그가 달린 논문 55편 · 필터 해제

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

2026-06-23 · Wonchul Shin, Inyong Choi, Kyogu Lee arxiv

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance …

Speech Extraction

IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments

2026-05-14 · Dinanath Padhya, Sajen Maharjan, Binita Adhikari, Ishwor Raj Pokharel arxiv

Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. W…

Speech Extraction

VorTEX: Various overlap ratio for Target speech EXtraction

2026-03-16 · Ro-hoon Oh, Jihwan Seol, Bugeun Kim arxiv

Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into be…

Speech Extraction

Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation

2025-12-07 · Jisoo Park, Seonghak Lee, Guisik Kim, Taewoo Kim 외 arxiv

Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivati…

Speech EnhancementSpeech ExtractionSpeech Separation

ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction

2025-11-09 · Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng 외 arxiv

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and p…

Speech Extraction

Neural Speech Extraction with Human Feedback

2025-08-05 · Malek Itani, Ashton Graves, Sefik Emre Eskimez, Shyamnath Gollakota arxiv

We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The ref…

Speech Extraction

TF-MLPNet: Tiny Real-Time Neural Speech Separation

2025-08-05 · Malek Itani, Tuochao Chen, Shyamnath Gollakota arxiv

Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelera…

Speech ExtractionSpeech Separation

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

2025-06-11 · Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng 외

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semanti…

Speech ExtractionTarget Speaker Extraction

Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction

2025-06-02 · Wang Dai, Archontis Politis, Tuomas Virtanen

We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relativ…

AttributeSpeech Extraction

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

2025-05-25 · Helin Wang, Jiarui Hai, Dongchao Yang, Chen Chen 외

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent a…

Speech ExtractionSpeech Separation

Single-Channel Target Speech Extraction Utilizing Distance and Room Clues

2025-05-20 · Runwu Shi, Zirui Lin, Benjamin Yen, Jiang Wang 외

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which c…

Speech ExtractionSpeech Separation

SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures

2025-04-15 · Kuang Yuan, Yifeng Wang, Xiyuxing Zhang, Chengyi Shen 외

Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce Sonic…

Speech Extraction

Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction

2025-03-11 · Minsu Kim, Rodrigo Mira, Honglie Chen, Stavros Petridis 외

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike …

Speech Extraction

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

2025-01-24 · Hao Ma, Rujin Chen, Xiao-Lei Zhang, Ju Liu 외

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spect…

Speech Extraction

Beyond Speaker Identity: Text Guided Target Speech Extraction

2025-01-15 · Mingyue Huo, Abhinav Jain, Cong Phuoc Huynh, Fanjie Kong 외

Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided…

Speech ExtractionSpeech Separation

Distance Based Single-Channel Target Speech Extraction

2024-12-28 · Runwu Shi, Benjamin Yen, Kazuhiro Nakadai

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological…

Speech Extraction

Investigation of Speaker Representation for Target-Speaker Speech Processing

2024-10-15 · Takanori Ashihara, Takafumi Moriya, Shota Horiguchi, Junyi Peng 외

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting infor…

Action DetectionActivity DetectionAutomatic Speech RecognitionSpeaker Recognition+4

Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration

2024-09-24 · Pin-Jui Ku, Alexander H. Liu, Roman Korostik, Sung-Feng Huang 외

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any v…

Bandwidth ExtensionDenoisingSpeech DenoisingSpeech Extraction+1

Look Once to Hear: Target Speech Hearing with Noisy Examples

2024-05-10 · Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka 외

In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achieves this capability, enabling target spe…

CPUSpeech Extraction

Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction

2024-04-19 · Zhaoxi Mu, Xinyu Yang

The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the chall…

Speech Extraction
1–20 / 55 다음 →