paper-with-me

홈 › Papers

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

2026-07-29 · Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, Chih-Pin Tan, Fang-Duo Tsai, Yi-Hsuan Yang arxiv

Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit linguistic information from V/UV tags cannot guarantee lyric accuracy, leading to a high phoneme error rate (PER). Inspired by singing voice synthesis (SVS), we propose MPEcho, which integrates a phoneme encoder and a length regulator (LR) into the SongEcho framework. By providing explicit phoneme-level conditioning and precise temporal boundaries, MPEcho significantly reduces PER. To enable this, we developed Phonsa, a Whisper-based automatic transcription model that provides high-precision phoneme-level annotations for singing voices, overcoming the scarcity of high-quality audio-phoneme pairs. Experimental results validate the effectiveness of Phonsa for alignment and MPEcho for end-to-end CSG. The audio samples, code and weights can be accessed from https://lonian6.github.io/MPEcho.github.io/.

📄 PDF Abstract BibTeX arXiv:2607.26698

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance

2025-12-04 · Junjie Zheng, Chunbo Hao, Guobin Ma, Xiaoyu Zhang 외 arxiv

Singing Voice Synthesis (SVS) remains constrained in practical deployment due to its strong dependence on accurate phoneme-level alignment and manually annotated melody contours, requirements that are resource-intensive …

Reinforcement Learning

A Melody-Unsupervision Model for Singing Voice Synthesis

2021-10-13 · Soonbeom Choi, Juhan Nam

Recent studies in singing voice synthesis have achieved high-quality results leveraging advances in text-to-speech models based on deep neural networks. One of the main issues in training singing voice synthesis models i…

modelSinging Voice Synthesistext-to-speechText to Speech

Speech-to-Singing Conversion in an Encoder-Decoder Framework

2020-02-16 · Jayneel Parekh, Preeti Rao, Yi-Hsuan Yang

In this paper our goal is to convert a set of spoken lines into sung ones. Unlike previous signal processing based methods, we take a learning based approach to the problem. This allows us to automatically model various …

DecoderMulti-Task Learning

Conditional LSTM-GAN for Melody Generation from Lyrics

2019-08-15 · Yi Yu, Abhishek Srivastava, Simon Canales

Melody generation from lyrics has been a challenging research issue in the field of artificial intelligence and music, which enables to learn and discover latent relationship between interesting lyrics and accompanying m…

Generative Adversarial Network

Melody transcription via generative pre-training

2022-12-04 · Chris Donahue, John Thickstun, Percy Liang

Despite the central role that melody plays in music perception, it remains an open challenge in music information retrieval to reliably detect the notes of the melody present in an arbitrary music recording. A key challe…

Chord RecognitionInformation RetrievalMusic Information RetrievalRetrieval