paper-with-me

홈 › Papers

DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

2025-08-08 · Wei Chen, Binzhu Sha, Dan Luo, Jing Yang, Zhuo Wang, Fan Fan, Zhiyong Wu arxiv

Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations.

📄 PDF Abstract BibTeX arXiv:2508.05978

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningAudio GenerationVoice Conversion

Similar Papers 제목 키워드 기반

Robust One-Shot Singing Voice Conversion

2022-10-20 · Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji

Recent progress in deep generative models has improved the quality of voice conversion in the speech domain. However, high-quality singing voice conversion (SVC) of unseen singers remains challenging due to the wider var…

Voice Conversion

Real-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis

2024-05-23 · Hui Li, Hongyu Wang, Zhijin Chen, Bohan Sun 외

Singing voice conversion is to convert the source singing voice into the target singing voice except for the content. Currently, flow-based models can complete the task of voice conversion, but they struggle to effective…

AttributeDecoderVoice Conversion

Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion

2024-06-04 · RuiQi Li, Rongjie Huang, Yongqi Wang, Zhiqing Hong 외

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal…

In-Context LearningLanguage ModelingLanguage ModellingRhythm+3

SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement

2024-07-10 · ZiHao Wang, Le Ma, Yongsheng Feng, Xin Pan 외

Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can hardly perform zero-shot due to incomplet…

DisentanglementVoice Conversion

LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance

2024-06-08 · Shihao Chen, Yu Gu, Jie Zhang, Na Li 외

Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few seconds of singing data. However, during the c…

Voice Conversion