paper-with-me

Papers Voice Cloning

“Voice Cloning” 태그가 달린 논문 112편 · 필터 해제

Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison

2025-07-15 · Andrew Valdivia, Yueming Zhang, Hailu Xu, Amir Ghasemkhani 외

This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions w…

Voice Cloning

De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks

2025-07-03 · Wei Fan, Kejiang Chen, Chang Liu, Weiming Zhang 외

The rapid advancement of speech generation models has heightened privacy and security concerns related to voice cloning (VC). Recent studies have investigated disrupting unauthorized voice cloning by introducing adversar…

Voice Cloning

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

2025-05-29 · Neta Glazer, David Chernin, Idan Achituve, Sharon Gannot 외

Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS systems continue to evolve, detection models m…

Audio Deepfake DetectionDeepFake DetectionFace SwappingGaussian Processes+3

Voice Adaptation for Swiss German

2025-05-28 · Samuel Stucki, Jan Deriu, Mark Cieliebak

This work investigates the performance of Voice Adaptation models for Swiss German dialects, i.e., translating Standard German text to Swiss German dialect speech. For this, we preprocess a large dataset of Swiss podcast…

Voice Cloning

VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents

2025-05-27 · Haiyun Li, Zhiyong Wu, XiaoFeng Xie, Jingran Xie 외

Voice cloning (VC)-resistant watermarking is an emerging technique for tracing and preventing unauthorized cloning. Existing methods effectively trace traditional VC models by training them on watermarked audio but fail …

Voice Cloning

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

2025-05-27 · Praveen Srinivasa Varadhan, Srija Anand, Soma Siddhartha, Mitesh M. Khapra

What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing…

Synthetic Data GenerationVoice Cloning

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

2025-05-25 · Renyuan Li, Zhibo Liang, Haichuan Zhang, Tianyu Shi 외

Recent breakthroughs in text-to-speech (TTS) voice cloning have raised serious privacy concerns, allowing highly accurate vocal identity replication from just a few seconds of reference audio, while retaining the speaker…

text-to-speechText to SpeechVoice Cloning

Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection

2025-05-22 · Jiaxin Liu, Jia Wang, Saihui Hou, Min Ren 외

In recent years, the rapid development of deepfake technology has given rise to an emerging and serious threat to public security: diffusion model-based digital human generation. Unlike traditional face manipulation meth…

DeepFake DetectionFace SwappingVoice Cloning

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

2025-05-21 · Cheng Yifan, Zhang Ruoyi, Shi Jiatong

Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional …

Emotion RecognitionFace DetectionLanguage ModelingLanguage Modelling+6

VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning

2025-05-18 · Qianyue Hu, Junyan Wu, Wei Lu, Xiangyang Luo

Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrup…

Representation LearningVoice Cloning

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

2025-05-12 · BoWen Zhang, Congchao Guo, Geng Yang, Hang Yu 외

We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a ref…

text-to-speechText to SpeechVoice Cloning

Voice Cloning: Comprehensive Survey

2025-05-01 · Hussam Azzuni, Abdulmotaleb El Saddik

Voice Cloning has rapidly advanced in today's digital world, with many researchers and corporations working to improve these algorithms for various applications. This article aims to establish a standardized terminology …

SurveyVoice Cloning

ClonEval: An Open Voice Cloning Benchmark

2025-04-29 · Iwona Christop, Tomasz Kuczyński, Marek Kubis

We present a novel benchmark for voice cloning text-to-speech models. The benchmark consists of an evaluation protocol, an open-source library for assessing the performance of voice cloning models, and an accompanying le…

text-to-speechText to SpeechVoice Cloning

"It's not a representation of me": Examining Accent Bias and Digital Exclusion in Synthetic AI Voice Services

2025-04-12 · Shira Michel, Sufi Kaur, Sarah Elizabeth Gillespie, Jeffrey Gleason 외

Recent advances in artificial intelligence (AI) speech generation and voice cloning technologies have produced naturalistic speech and accurate voice replication, yet their influence on sociotechnical systems across dive…

Voice Cloning

Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis

2025-04-10 · Yizhong Geng, Jizhuo Xu, Zeyu Liang, Jinghan Yang 외

Text-to-speech (TTS) technology has achieved impressive results for widely spoken languages, yet many under-resourced languages remain challenged by limited data and linguistic complexities. In this paper, we present a n…

Speech Synthesistext-to-speechText to SpeechVoice Cloning

SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development

2025-03-31 · Minghan Wang, Ye Bai, Yuxia Wang, Thuy-Trang Vu 외

High-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations. Human recordings incur high costs and privacy concerns, while synthetic approac…

Speech SynthesisVoice Cloning

SoK: How Robust is Audio Watermarking in Generative AI models?

2025-03-24 · Yizhu Wen, Ashwin Innuganti, Aaron Bien Ramos, Hanqing Guo 외

Audio watermarking is increasingly used to verify the provenance of AI-generated content, enabling applications such as detecting AI-generated speech, protecting music IP, and defending against voice cloning. To be effec…

Voice Cloning

Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology

2025-03-03 · Birger Moell, Fredrik Sand Aronsson

This study explores voice cloning to generate synthetic speech replicating the unique patterns of individuals with dysarthria. Using the TORGO dataset, we address data scarcity and privacy challenges in speech-language p…

Speech SynthesisVoice Cloning

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

2025-03-03 · Xinsheng Wang, Mingqi Jiang, Ziyang Ma, Ziyu Zhang 외

Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis. However, existing foundation models rely on multi-stage processing or complex architecture…

Attributetext-to-speechText to SpeechVoice Cloning

Steganography Beyond Space-Time with Chain of Multimodal AI

2025-02-25 · Ching-Chun Chang, Isao Echizen

Steganography is the art and science of covert writing, with a broad range of applications interwoven within the realm of cybersecurity. As artificial intelligence continues to evolve, its ability to synthesise realistic…

Face SwappingText GenerationVoice Cloning
1–20 / 112 다음 →