paper-with-me

Papers

Acoustic Modeling for Automatic Lyrics-to-Audio Alignment

2019-06-25 · Chitralekha Gupta, Emre Yilmaz, Haizhou Li

Automatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we propose (1) using additional speech and music-informed features and (2) adapting the acoustic models trained on a large amount of solo singing vocals towards polyphonic music using a small amount of in-domain data. Incorporating additional information such as voicing and auditory features together with conventional acoustic features aims to bring robustness against the increased spectro-temporal variations in singing vocals. By adapting the acoustic model using a small amount of polyphonic audio data, we reduce the domain mismatch between training and testing data. We perform several alignment experiments and present an in-depth alignment error analysis on acoustic features, and model adaptation techniques. The results demonstrate that the proposed strategy provides a significant error reduction of word boundary alignment over comparable existing systems, especially on more challenging polyphonic data with long-duration musical interludes.

📄 PDF Abstract BibTeX arXiv:1906.10369

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On-Line Audio-to-Lyrics Alignment Based on a Reference Performance

2021-07-30 · Charles Brazier, Gerhard Widmer

Audio-to-lyrics alignment has become an increasingly active research task in MIR, supported by the emergence of several open-source datasets of audio recordings with word-level lyrics annotations. However, there are stil…

Specificity

Lyrics-to-Audio Alignment by Unsupervised Discovery of Repetitive Patterns in Vowel Acoustics

2017-01-21 · Sungkyun Chang, Kyogu Lee

Most of the previous approaches to lyrics-to-audio alignment used a pre-developed automatic speech recognition (ASR) system that innately suffered from several difficulties to adapt the speech model to individual singers…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

DeepSinger: Singing Voice Synthesis with Data Mined From the Web

2020-07-09 · Yi Ren, Xu Tan, Tao Qin, Jian Luan 외

In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites. The pipeline of DeepSinger consi…

SentenceSinging Voice Synthesis

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

2025-07-28 · Renhang Liu, Chia-Yu Hung, Navonil Majumder, Taylor Gautreaux 외 arxiv

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech a…

Audio Generation

MSTRE-Net: Multistreaming Acoustic Modeling for Automatic Lyrics Transcription

2021-08-05 · Emir Demirel, Sven Ahlbäck, Simon Dixon

This paper makes several contributions to automatic lyrics transcription (ALT) research. Our main contribution is a novel variant of the Multistreaming Time-Delay Neural Network (MTDNN) architecture, called MSTRE-Net, wh…

Automatic Lyrics TranscriptionTAG