paper-with-me

홈 › Papers

Tutti: Expressive Multi-Singer Synthesis via Structure-Level Timbre Control and Vocal Texture Modeling

2026-02-09 · Jiatao Chen, Xing Tang, Xiaoyue Duan, Yutang Feng, Jinchao Zhang, Jie Zhou arxiv

While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address this, we propose Tutti, a unified framework designed for structured multi-singer generation. Specifically, we introduce a Structure-Aware Singer Prompt to enable flexible singer scheduling evolving with musical structure, and propose Complementary Texture Learning via Condition-Guided VAE to capture implicit acoustic textures (e.g., spatial reverberation and spectral fusion) that are complementary to explicit controls. Experiments demonstrate that Tutti excels in precise multi-singer scheduling and significantly enhances the acoustic realism of choral generation, offering a novel paradigm for complex multi-singer arrangement. Audio samples are available at https://annoauth123-ctrl.github.io/Tutii_Demo/.

📄 PDF Abstract BibTeX arXiv:2602.08233

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

2025-09-24 · Junchuan Zhao, Wei Zeng, Tianle Lyu, Ye Wang arxiv

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled z…

Contrastive LearningSpeech Synthesis

SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis

2026-02-08 · Jiale Qian, Hao Meng, Tian Zheng, Pengcheng Zhu 외 arxiv

While recent years have witnessed rapid progress in speech synthesis, open-source singing voice synthesis (SVS) systems still face significant barriers to industrial deployment, particularly in terms of robustness and ze…

Zero-shot GeneralizationSpeech Synthesis

Towards Improving the Expressiveness of Singing Voice Synthesis with BERT Derived Semantic Information

2023-08-31 · Shaohuan Zhou, Shun Lei, Weiya You, Deyi Tuo 외

This paper presents an end-to-end high-quality singing voice synthesis (SVS) system that uses bidirectional encoder representation from Transformers (BERT) derived semantic embeddings to improve the expressiveness of the…

Singing Voice Synthesis

SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture

2025-06-26 · Kehan Sui, Jinxu Xiang, Fang Jin

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkabl…

DenoisingSinging Voice SynthesisVideo Generation

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

2021-05-06 · Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 외

Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous s…

Generative Adversarial NetworkSinging Voice Synthesistext-to-speechText to Speech+1