paper-with-me

Papers

Controllable and Interpretable Singing Voice Decomposition via Assem-VC

2021-10-25 · Kang-wook Kim, Junhyeok Lee

We propose a singing decomposition system that encodes time-aligned linguistic content, pitch, and source speaker identity via Assem-VC. With decomposed speaker-independent information and the target speaker's embedding, we could synthesize the singing voice of the target speaker. In conclusion, we made a perfectly synced duet with the user's singing voice and the target singer's converted singing voice.

📄 PDF Abstract BibTeX arXiv:2110.12676

Code (1)

mindslab-ai/assem-vc 공식 구현 pytorch

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation

2025-08-22 · Xueyao Zhang, Junan Zhang, Yuancheng Wang, Chaoren Wang 외 arxiv

Controllable human voice generation, particularly for expressive domains like singing, remains a significant challenge. This paper introduces Vevo2, a unified framework for controllable speech and singing voice generatio…

VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion

2025-05-27 · Joon-Seung Choi, Dong-Min Byun, Hyung-Seok Oh, Seong-Whan Lee

Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling v…

Voice Conversion

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

2024-09-20 · Yu Zhang, Changhao Pan, Wenxiang Guo, RuiQi Li 외

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited div…

AllSinging Voice SynthesisStyle TransferVocal technique classification

M4Singer: a Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

2022-12-29 · NIPS 2022 12 · Lichao Zhang, RuiQi Li, Shoutong Wang, Liqun Deng 외

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Mult…

Music TranscriptionSinging Voice SynthesisVoice Conversion

UniVoice: A Unified Model for Speech and Singing Voice Generation

2026-06-04 · Junjie Zheng, Huixin Xue, Shihong Ren, Chaofan Ding 외 arxiv

Text-to-speech (TTS) and singing voice synthesis (SVS) both aim to generate human vocal audio from symbolic inputs, but they impose different requirements on the generation process. Speech generation relies on flexible, …