paper-with-me

홈 › Papers

Robust Singing Voice Transcription Serves Synthesis

2024-05-16 · RuiQi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong, Rongjie Huang, Zhou Zhao

Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AST methods, however, struggle with accuracy and robustness when used for practical annotation. This paper presents ROSVOT, the first robust AST model that serves SVS, incorporating a multi-scale framework that effectively captures coarse-grained note information and ensures fine-grained frame-level segmentation, coupled with an attention-based pitch decoder for reliable pitch prediction. We also established a comprehensive annotation-and-training pipeline for SVS to test the model in real-world settings. Experimental findings reveal that ROSVOT achieves state-of-the-art transcription accuracy with either clean or noisy inputs. Moreover, when trained on enlarged, automatically annotated datasets, the SVS model outperforms its baseline, affirming the capability for practical application. Audio samples are available at https://rosvot.github.io.

📄 PDF Abstract BibTeX arXiv:2405.09940

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSinging Voice Synthesis

Similar Papers 제목 키워드 기반

A Singing Voice Database in Basque for Statistical Singing Synthesis of Bertsolaritza

2016-05-01 · LREC 2016 5 · Xabier Sarasola, Eva Navas, David Tavarez, Daniel Erro 외

This paper describes the characteristics and structure of a Basque singing voice database of bertsolaritza. Bertsolaritza is a popular singing style from Basque Country sung exclusively in Basque that is improvised and a…

Singing Voice Synthesis

End-to-end lyrics Recognition with Voice to Singing Style Transfer

2021-02-17 · Sakya Basak, Shrutina Agarwal, Sriram Ganapathy, Naoya Takahashi

Automatic transcription of monophonic/polyphonic music is a challenging task due to the lack of availability of large amounts of transcribed data. In this paper, we propose a data augmentation method that converts natura…

Data AugmentationLanguage ModelingLanguage ModellingStyle Transfer+1

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models

2026-05-06 · Yukun Chen, Tianrui Wang, Zhaoxi Mu, Xinyu Yang 외 arxiv

High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musi…

M4Singer: a Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

2022-12-29 · NIPS 2022 12 · Lichao Zhang, RuiQi Li, Shoutong Wang, Liqun Deng 외

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Mult…

Music TranscriptionSinging Voice SynthesisVoice Conversion

Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss

2025-06-03 · Jiawen Huang, Felipe Sousa, Emir Demirel, Emmanouil Benetos 외

Automatic Lyrics Transcription (ALT) aims to recognize lyrics from singing voices, similar to Automatic Speech Recognition (ASR) for spoken language, but faces added complexity due to domain-specific properties of the si…

Automatic Lyrics TranscriptionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1