Beyond Speaker Identity: Text Guided Target Speech Extraction
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in addition to the audio clue to extract the desired speech from a given mixture. Our model integrates a speech separation network adapted from SepFormer with a bi-modality clue network that flexibly processes both audio and text clues. To train and evaluate our model, we introduce a new dataset TextrolMix with speech mixtures and natural language descriptions. Experimental results demonstrate that our method effectively separates speech based not only on who is speaking, but also on how they are speaking, enhancing TSE in scenarios where traditional audio clues are absent. Demos are at: https://mingyue66.github.io/TextrolMix/demo/
Code (1)
Tasks
Speech ExtractionSpeech SeparationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Evaluating X-vector-based Speaker Anonymization under White-box Assessment
In the scenario of the Voice Privacy challenge, anonymization is achieved by converting all utterances from a source speaker to match the same target identity; this identity being randomly selected. In this context, an a…
Speaker anonymizationGuided Training: A Simple Method for Single-channel Speaker Separation
Deep learning has shown a great potential for speech separation, especially for speech and non-speech separation. However, it encounters permutation problem for multi-speaker separation where both target and interference…
Speaker SeparationSpeech SeparationStyleTTS-VC: One-Shot Voice Conversion by Knowledge Transfer from Style-Based TTS Models
One-shot voice conversion (VC) aims to convert speech from any source speaker to an arbitrary target speaker with only a few seconds of reference speech from the target speaker. This relies heavily on disentangling the s…
Data Augmentationtext-to-speechText to SpeechTransfer Learning+1VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrup…
Representation LearningVoice CloningCross-speaker style transfer for text-to-speech using data augmentation
We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting…
Data AugmentationStyle Transfertext-to-speechText to Speech+1