paper-with-me

Papers Target Speaker Extraction

“Target Speaker Extraction” 태그가 달린 논문 55편 · 필터 해제

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

2025-06-11 · Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng 외

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semanti…

Speech ExtractionTarget Speaker Extraction

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

2025-05-31 · Cunhang Fan, Ying Chen, Jian Zhou, Zexu Pan 외

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlo…

Contrastive LearningEEGTarget Speaker Extraction

FlowTSE: Target Speaker Extraction with Flow Matching

2025-05-20 · Aviv Navon, Aviv Shamsian, Yael Segal-Feldman, Neta Glazer 외

Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE ach…

Target Speaker Extraction

Listen to Extract: Onset-Prompted Target Speaker Extraction

2025-05-08 · Pengjie Shen, Kangrui Chen, Shulin He, Pengru Chen 외

We propose $\textit{listen to extract}$ (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting…

Target Speaker Extraction

LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models

2025-04-10 · Beilong Tang, Bang Zeng, Ming Li

We propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction (TSE) based on the LauraGPT backbone. It employs a small-scale auto-regressive decoder-only language model which takes the…

DecoderLanguage ModelingLanguage ModellingTarget Speaker Extraction

$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction

2025-04-01 · Wenxuan Wu, Xueyuan Chen, Shuai Wang, Jiadong Wang 외

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals…

Target Speaker Extraction

Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments

2025-02-23 · Shitong Xu, Yiyuan Yang, Niki Trigoni, Andrew Markham

Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean au…

Target Speaker Extraction

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

2025-02-05 · Yuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang 외

We introduce Metis, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled …

Self-Supervised LearningSpeech EnhancementTarget Speaker Extractiontext-to-speech+2

AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement

2025-01-26 · Junan Zhang, Jing Yang, Zihao Fang, Yuancheng Wang 외

We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable of handling both speech and singing voice…

DenoisingIn-Context LearningSuper-ResolutionTarget Speaker Extraction

Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection

2025-01-07 · Bang Zeng, Ming Li

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction…

Action DetectionActivity DetectionAutomatic Speech RecognitionMulti-Task Learning+5

MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

2024-12-11 · Junjie Li, Ke Zhang, Shuai Wang, Kong Aik Lee 외

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always avail…

Target Speaker Extraction

Multi-Level Speaker Representation for Target Speaker Extraction

2024-10-21 · Ke Zhang, Junjie Li, Shuai Wang, Yangjie Wei 외

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with…

Target Speaker Extraction

STCON System for the CHiME-8 Challenge

2024-10-17 · Anton Mitrofanov, Tatiana Prisyach, Tatiana Timofeeva, Sergei Novoselov 외

This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trai…

Data AugmentationSpeech SeparationTarget Speaker Extraction

Wanna hear your voice? A sample is all we need!

2024-10-01 · The Hieu Pham, Phuong Thanh Tran Nguyen, Xuan Tho Nguyen, Tan Dat Nguyen 외

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain u…

AllSpeech SeparationTarget Speaker Extraction

Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions

2024-09-29 · Jinyi Mi, Xiaohan Shi, Ding Ma, Jiajun He 외

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limit…

Emotion RecognitionSpeech Emotion RecognitionTarget Speaker Extraction

Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration

2024-09-24 · Pin-Jui Ku, Alexander H. Liu, Roman Korostik, Sung-Feng Huang 외

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any v…

Bandwidth ExtensionDenoisingSpeech DenoisingSpeech Extraction+1

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

2024-09-24 · Shuai Wang, Ke Zhang, Shaoxiong Lin, Junjie Li 외

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increas…

Managementspeech-recognitionSpeech RecognitionTarget Speaker Extraction

TSELM: Target Speaker Extraction using Discrete Tokens and Language Models

2024-09-12 · Beilong Tang, Bang Zeng, Ming Li

We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mecha…

Audio GenerationTarget Speaker Extraction

USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

2024-09-04 · Bang Zeng, Ming Li

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recogniti…

Speaker RecognitionSpeech SeparationTarget Speaker Extraction

Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement

2024-09-02 · Tathagata Bandyopadhyay

Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a trans…

Target Speaker Extraction
1–20 / 55 다음 →