TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate target speaker information. Language models are employed to capture the sequence dependencies, while a scalable HiFi-GAN is used to reconstruct the audio from the tokens. By applying a cross-entropy loss, TSELM models the probability distribution of output tokens, thus converting the complex regression problem of audio generation into a classification task. Experimental results show that TSELM achieves excellent results in speech quality and comparable results in speech intelligibility.
Code (1)
Tasks
Audio GenerationTarget Speaker ExtractionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic…
In-Context LearningVoice ConversionInter-Speaker Relative Cues for Text-Guided Target Speech Extraction
We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relativ…
AttributeSpeech ExtractionSpeaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation…
Boundary DetectionHigh Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting mod…
Computational EfficiencyLanguage ModelingLanguage ModellingSpeech Synthesis+2LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models
We propose LauraTSE, an Auto-Regressive Decoder-Only Language Model for Target Speaker Extraction (TSE) based on the LauraGPT backbone. It employs a small-scale auto-regressive decoder-only language model which takes the…
DecoderLanguage ModelingLanguage ModellingTarget Speaker Extraction