paper-with-me

Papers

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck

2026-04-07 · Zhetao Hu, Yiquan Zhou, Wenyu Wang, Zhiyu Wu, Xin Gao, Jihua Zhu arxiv

This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)-a novel singing style conversion system that advances fine-grained style conversion and control within in-domain settings. To address the critical challenges of style leakage, dynamic rendering, and high-fidelity generation with limited data, we introduce three key innovations: a boundary-aware Whisper bottleneck that pools phoneme-span representations to suppress residual source style while preserving linguistic content; an explicit frame-level technique matrix, enhanced by targeted F0 processing during inference, for stable and distinct dynamic style rendering; and a perceptually motivated high-frequency band completion strategy that leverages an auxiliary standard 48kHz SVC model to augment the high-frequency spectrum, thereby overcoming data scarcity without overfitting. In the official SVCC2025 subjective evaluation, our system achieves the best naturalness performance among all submissions while maintaining competitive results in speaker similarity and technique control, despite using significantly less extra singing data than other top-performing systems. Audio samples are available online.

📄 PDF Abstract BibTeX arXiv:2604.05526

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion

2025-05-27 · Joon-Seung Choi, Dong-Min Byun, Hyung-Seok Oh, Seong-Whan Lee

Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling v…

Voice Conversion

Speech-to-Singing Conversion based on Boundary Equilibrium GAN

2020-05-28 · Da-Yi Wu, Yi-Hsuan Yang

This paper investigates the use of generative adversarial network (GAN)-based models for converting the spectrogram of a speech signal into that of a singing one, without reference to the phoneme sequence underlying the …

DecoderGenerative Adversarial NetworkStyle Transfer

Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation

2025-08-22 · Xueyao Zhang, Junan Zhang, Yuancheng Wang, Chaoren Wang 외 arxiv

Controllable human voice generation, particularly for expressive domains like singing, remains a significant challenge. This paper introduces Vevo2, a unified framework for controllable speech and singing voice generatio…

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

2024-09-20 · Yu Zhang, Changhao Pan, Wenxiang Guo, RuiQi Li 외

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited div…

AllSinging Voice SynthesisStyle TransferVocal technique classification

Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control

2026-06-15 · Joon-Seung Choi, Dong-Min Byun, Seong-Whan Lee arxiv

Singing style is a crucial aspect of a natural and expressive singing voice. Singers utilize singing styles to convey the feeling or emotion of the songs. Several works have been proposed to control singing style for mak…

Voice Conversion