paper-with-me

홈 › Papers

LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning

2025-07-07 · Sandipan Dhar, Mayank Gupta, Preeti Rao arxiv

The field of Singing Voice Synthesis (SVS) has seen significant advancements in recent years due to the rapid progress of diffusion-based approaches. However, capturing vocal style, genre-specific pitch inflections, and language-dependent characteristics remains challenging, particularly in low-resource scenarios. To address this, we propose LAPS-Diff, a diffusion model integrated with language-aware embeddings and a vocal-style guided learning mechanism, specifically designed for Bollywood Hindi singing style. We curate a Hindi SVS dataset and leverage pre-trained language models to extract word and phone-level embeddings for an enriched lyrics representation. Additionally, we incorporated a style encoder and a pitch extraction model to compute style and pitch losses, capturing features essential to the naturalness and expressiveness of the synthesized singing, particularly in terms of vocal style and pitch variations. Furthermore, we utilize MERT and IndicWav2Vec models to extract musical and contextual embeddings, serving as conditional priors to refine the acoustic feature generation process further. Based on objective and subjective evaluations, we demonstrate that LAPS-Diff significantly improves the quality of the generated samples compared to the considered state-of-the-art (SOTA) model for our constrained dataset that is typical of the low resource scenario.

📄 PDF Abstract BibTeX arXiv:2507.04966

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HiddenSinger: High-Quality Singing Voice Synthesis via Neural Audio Codec and Latent Diffusion Models

2023-06-12 · Ji-Sang Hwang, Sang-Hoon Lee, Seong-Whan Lee

Recently, denoising diffusion models have demonstrated remarkable performance among generative models in various domains. However, in the speech domain, the application of diffusion models for synthesizing time-varying a…

DenoisingSinging Voice SynthesisSpeech Synthesis

MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance

2024-06-10 · Semin Kim, Myeonghun Jeong, Hyeonseung Lee, Minchan Kim 외

In this paper, we propose MakeSinger, a semi-supervised training method for singing voice synthesis (SVS) via classifier-free diffusion guidance. The challenge in SVS lies in the costly process of gathering aligned sets …

Singing Voice Synthesistext-to-speechText to Speech

LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance

2024-06-08 · Shihao Chen, Yu Gu, Jie Zhang, Na Li 외

Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few seconds of singing data. However, during the c…

Voice Conversion

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

2021-05-06 · Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 외

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous s…

Generative Adversarial NetworkSinging Voice Synthesistext-to-speechText to Speech+1

Hierarchical Diffusion Models for Singing Voice Neural Vocoder

2022-10-14 · Naoya Takahashi, Mayank Kumar, Singh, Yuki Mitsufuji

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions i…