Singer Identity Representation Learning using Self-Supervised Techniques
Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework for training singer identity encoders to extract representations suitable for various singing-related tasks, such as singing voice similarity and synthesis. We explore different self-supervised learning techniques on a large collection of isolated vocal tracks and apply data augmentations during training to ensure that the representations are invariant to pitch and content variations. We evaluate the quality of the resulting representations on singer similarity and identification tasks across multiple datasets, with a particular emphasis on out-of-domain generalization. Our proposed framework produces high-quality embeddings that outperform both speaker verification and wav2vec 2.0 pre-trained baselines on singing voice while operating at 44.1 kHz. We release our code and trained models to facilitate further research on singing voice and related areas.
Code (1)
Tasks
Domain GeneralizationRepresentation LearningSelf-Supervised LearningSpeaker VerificationVoice SimilaritySimilar Papers 제목 키워드 기반
Self-Supervised Representations for Singing Voice Conversion
A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Ve…
DisentanglementVoice ConversionSinging Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders
We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages re…
DecoderVoice ConversionSelf-Supervised Contrastive Learning for Singing Voices
This study introduces self-supervised contrastive learning to acquire feature representations of singing voices. To acquire robust representations in an unsupervised manner, regular self-supervised contrastive learning t…
Contrastive LearningSinger IdentificationVocal technique classificationMulti-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus
High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not me…
Audio GenerationSinging Voice SynthesisText-To-Speech SynthesisVAW-GAN for Singing Voice Conversion with Non-parallel Training Data
Singing voice conversion aims to convert singer's voice from source to target without changing singing content. Parallel training data is typically required for the training of singing voice conversion system, that is ho…
DecoderGenerative Adversarial NetworkVoice Conversion