Papers Target Speaker Extraction
“Target Speaker Extraction” 태그가 달린 논문 55편 · 필터 해제
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semanti…
Speech ExtractionTarget Speaker ExtractionM3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlo…
Contrastive LearningEEGTarget Speaker ExtractionFlowTSE: Target Speaker Extraction with Flow Matching
Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE ach…
Target Speaker ExtractionListen to Extract: Onset-Prompted Target Speaker Extraction
We propose $\textit{listen to extract}$ (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting…
Target Speaker ExtractionLauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models
We propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction (TSE) based on the LauraGPT backbone. It employs a small-scale auto-regressive decoder-only language model which takes the…
DecoderLanguage ModelingLanguage ModellingTarget Speaker Extraction$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals…
Target Speaker ExtractionTarget Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean au…
Target Speaker ExtractionMetis: A Foundation Speech Generation Model with Masked Generative Pre-training
We introduce Metis, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled …
Self-Supervised LearningSpeech EnhancementTarget Speaker Extractiontext-to-speech+2AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable of handling both speech and singing voice…
DenoisingIn-Context LearningSuper-ResolutionTarget Speaker ExtractionUniversal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction…
Action DetectionActivity DetectionAutomatic Speech RecognitionMulti-Task Learning+5MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always avail…
Target Speaker ExtractionMulti-Level Speaker Representation for Target Speaker Extraction
Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with…
Target Speaker ExtractionSTCON System for the CHiME-8 Challenge
This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trai…
Data AugmentationSpeech SeparationTarget Speaker ExtractionWanna hear your voice? A sample is all we need!
Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain u…
AllSpeech SeparationTarget Speaker ExtractionTwo-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limit…
Emotion RecognitionSpeech Emotion RecognitionTarget Speaker ExtractionGenerative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any v…
Bandwidth ExtensionDenoisingSpeech DenoisingSpeech Extraction+1WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increas…
Managementspeech-recognitionSpeech RecognitionTarget Speaker ExtractionTSELM: Target Speaker Extraction using Discrete Tokens and Language Models
We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mecha…
Audio GenerationTarget Speaker ExtractionUSEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recogniti…
Speaker RecognitionSpeech SeparationTarget Speaker ExtractionSpectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a trans…
Target Speaker Extraction