paper-with-me

홈 › Papers

Adversarially Trained Multi-Singer Sequence-To-Sequence Singing Synthesizer

2020-06-18 · Jie Wu, Jian Luan

This paper presents a high quality singing synthesizer that is able to model a voice with limited available recordings. Based on the sequence-to-sequence singing model, we design a multi-singer framework to leverage all the existing singing data of different singers. To attenuate the issue of musical score unbalance among singers, we incorporate an adversarial task of singer classification to make encoder output less singer dependent. Furthermore, we apply multiple random window discriminators (MRWDs) on the generated acoustic features to make the network be a GAN. Both objective and subjective evaluations indicate that the proposed synthesizer can generate higher quality singing voice than baseline (4.12 vs 3.53 in MOS). Especially, the articulation of high-pitched vowels is significantly enhanced.

📄 PDF Abstract BibTeX arXiv:2006.10317

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE using Mel-spectrograms

2021-10-08 · Chien-Feng Liao, Jen-Yu Liu, Yi-Hsuan Yang

In this paper, we propose a novel neural network model called KaraSinger for a less-studied singing voice synthesis (SVS) task named score-free SVS, in which the prosody and melody are spontaneously decided by machine. K…

Language ModelingLanguage ModellingSinging Voice Synthesis

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

2025-09-24 · Junchuan Zhao, Wei Zeng, Tianle Lyu, Ye Wang arxiv

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled z…

Contrastive LearningSpeech Synthesis

Tutti: Expressive Multi-Singer Synthesis via Structure-Level Timbre Control and Vocal Texture Modeling

2026-02-09 · Jiatao Chen, Xing Tang, Xiaoyue Duan, Yutang Feng 외 arxiv

While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single…

Future Frame Prediction of a Video Sequence

2020-08-31 · Jasmeen Kaur, Sukhendu Das

Predicting future frames of a video sequence has been a problem of high interest in the field of Computer Vision as it caters to a multitude of applications. The ability to predict, anticipate and reason about future eve…

Autonomous DrivingDecision MakingPredictionRobot Navigation