Voice Conversion
3개 벤치마크 · 논문 577편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks
AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization
Utilizing Self-supervised Representations for MOS Prediction
MOSNet: Deep Learning based Objective Assessment for Voice Conversion
Papers
Your Voice Cloning System is Secretly a Voice Anonymizer
Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speec…
Voice ConversionREIMU: Efficient Heterogeneous Hierarchical Reasoning for SSL-Based Speech Deepfake Detection
The increasing realism of speech generated by text-to-speech and voice conversion systems poses growing challenges to media integrity and voice authentication. Self-supervised learning (SSL) has substantially advanced sp…
Self-Supervised LearningDeepFake DetectionVoice ConversionTeffic-Audio: Tell Fact from Fiction
Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoof…
DeepFake DetectionSpeech SynthesisVoice ConversionProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed spe…
Voice ConversionNouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS). Existing SAS approaches modify voice …
Speaker VerificationVoice ConversionSpeech SynthesisGRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but inherit the ambiguity of text and mispr…
Voice Conversion