paper-with-me

Papers

VorTEX: Various overlap ratio for Target speech EXtraction

2026-03-16 · Ro-hoon Oh, Jihwan Seol, Bugeun Kim arxiv

Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into behavior across realistic overlap ratios. We introduce VorTEX (Various overlap ratio for Target speech EXtraction), a text-prompted TSE architecture with a Decoupled Adaptive Multi-branch (DAM) Fusion block that separates primary extraction from auxiliary regularization pathways. To enable controlled analysis, we construct PORTE, a two-speaker dataset spanning overlap ratios from 0% to 100%. We further propose Suppression Ratio on Energy (SuRE), a diagnostic metric that detects suppression behavior not captured by conventional measures. Experiments show that existing models exhibit suppression or residual interference under overlap, whereas VorTEX achieves the highest separation fidelity across 20-100% overlap (e.g., 5.50 dB at 20% and 2.04 dB at 100%) while maintaining zero SuRE, indicating robust extraction without suppression-driven artifacts.

📄 PDF Abstract BibTeX arXiv:2603.14803

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Extraction

Similar Papers 제목 키워드 기반

Low-Interference Near-Field Multi-User Communication Enabled by Spatially Converging Multi-Mode Vortex Waves

2025-02-18 · Yufei Zhao, Qihao Lv, Yuanbin Chen, Afkar Mohamed Ismail 외

This paper proposes a multi-user Spatial Division Multiplexing (SDM) near-field access scheme, inspired by the orthogonal characteristics of multi-mode vortex waves. A Reconfigurable Meta-surface (RM) is ingeniously empl…

Sparsely Overlapped Speech Training in the Time Domain: Joint Learning of Target Speech Separation and Personal VAD Benefits

2021-06-28 · Qingjian Lin, Lin Yang, Xuyang Wang, Luyuan Xie 외

Target speech separation is the process of filtering a certain speaker's voice out of speech mixtures according to the additional speaker identity information provided. Recent works have made considerable improvement by …

Speech Separation

X-VORTEX: Spatio-Temporal Contrastive Learning for Wake Vortex Trajectory Forecasting

2026-02-13 · Zhan Qu, Michael Färber arxiv

Wake vortices are strong, coherent air turbulences created by aircraft, and they pose a major safety and capacity challenge for air traffic management. Tracking how vortices move, weaken, and dissipate over time from LiD…

Trajectory ForecastingContrastive Learning

VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition

2022-02-22 · Jinhan Wang, Xiaosu Tong, Jinxi Guo, Di He 외

While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed methods, (partial) overlapping inference…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+3

USEV: Universal Speaker Extraction with Visual Cue

2021-09-30 · Zexu Pan, Meng Ge, Haizhou Li

A speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. H…