paper-with-me

Voice Conversion

3개 벤치마크 · 논문 577편 · 이 태스크의 논문 보기 →

Benchmarks

VCTK

결과 3개

Most implemented

Papers

Your Voice Cloning System is Secretly a Voice Anonymizer

2026-08-27 · Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu arxiv

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speec…

Voice Conversion

REIMU: Efficient Heterogeneous Hierarchical Reasoning for SSL-Based Speech Deepfake Detection

2026-08-01 · Kwok-Ho Ng, Tingting Song, Bingwen Feng, Peiya Li arxiv

The increasing realism of speech generated by text-to-speech and voice conversion systems poses growing challenges to media integrity and voice authentication. Self-supervised learning (SSL) has substantially advanced sp…

Self-Supervised LearningDeepFake DetectionVoice Conversion

Teffic-Audio: Tell Fact from Fiction

2026-07-30 · Wan Lin, Li Wang, Jindong Wang, Kunyu Feng 외 arxiv

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoof…

DeepFake DetectionSpeech SynthesisVoice Conversion

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

2026-07-06 · Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez, Jesus Villalba Lopez 외 arxiv

Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed spe…

Voice Conversion

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

2026-07-04 · Meiying Melissa Chen, Anastasia Kuznetsova, Zhenyu Wang, Zhiyao Duan arxiv

Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS). Existing SAS approaches modify voice …

Speaker VerificationVoice ConversionSpeech Synthesis

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

2026-07-02 · Antonis Asonitis, Francesco Verdini, Aref Farhadipour, Vijeta Avijeet 외 arxiv

We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but inherit the ambiguity of text and mispr…

Voice Conversion

전체 577편 보기 →