paper-with-me

Papers

FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

2024-05-13 · Jianyi Chen, Wei Xue, Xu Tan, Zhen Ye, Qifeng Liu, Yike Guo

Singing Accompaniment Generation (SAG), which generates instrumental music to accompany input vocals, is crucial to developing human-AI symbiotic art creation systems. The state-of-the-art method, SingSong, utilizes a multi-stage autoregressive (AR) model for SAG, however, this method is extremely slow as it generates semantic and acoustic tokens recursively, and this makes it impossible for real-time applications. In this paper, we aim to develop a Fast SAG method that can create high-quality and coherent accompaniments. A non-AR diffusion-based framework is developed, which by carefully designing the conditions inferred from the vocal signals, generates the Mel spectrogram of the target accompaniment directly. With diffusion and Mel spectrogram modeling, the proposed method significantly simplifies the AR token-based SingSong framework, and largely accelerates the generation. We also design semantic projection, prior projection blocks as well as a set of loss functions, to ensure the generated accompaniment has semantic and rhythm coherence with the vocal signal. By intensive experimental studies, we demonstrate that the proposed method can generate better samples than SingSong, and accelerate the generation by at least 30 times. Audio samples and code are available at https://fastsag.github.io/.

📄 PDF Abstract BibTeX arXiv:2405.07682

Code (0)

등록된 구현이 없습니다.

Tasks

Rhythm

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
SAG 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment

2024-04-14 · Zhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang 외

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, …

Music GenerationSinging Voice Synthesis

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

2026-06-05 · Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang, Yuxin Guo 외 arxiv

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment syne…

Voice Conversion

Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations

2024-11-03 · Quoc-Huy Trinh, Minh-Van Nguyen, Trong-Hieu Nguyen Mau, Khoa Tran 외

Singing is one of the most cherished forms of human entertainment. However, creating a beautiful song requires an accompaniment that complements the vocals and aligns well with the song instruments and genre. With advanc…

Investigation of Singing Voice Separation for Singing Voice Detection in Polyphonic Music

2020-04-08 · Yifu Sun, xulong Zhang, Yi Yu, Xi Chen 외

Singing voice detection (SVD), to recognize vocal parts in the song, is an essential task in music information retrieval (MIR). The task remains challenging since singing voice varies and intertwines with the accompanime…

Information RetrievalMelody ExtractionMusic Information RetrievalRetrieval

MLP Singer: Towards Rapid Parallel Singing Voice Synthesis

2021-06-15 · arXiv 2021 6 · Jaesung Tae, Hyeongju Kim, Younggun Lee

Recent developments in deep learning have significantly improved the quality of synthesized singing voice audio. However, prominent neural singing voice synthesis systems suffer from slow inference speed due to their aut…

image-classificationSinging Voice Synthesis