paper-with-me

Papers

On-Line Audio-to-Lyrics Alignment Based on a Reference Performance

2021-07-30 · Charles Brazier, Gerhard Widmer

Audio-to-lyrics alignment has become an increasingly active research task in MIR, supported by the emergence of several open-source datasets of audio recordings with word-level lyrics annotations. However, there are still a number of open problems, such as a lack of robustness in the face of severe duration mismatches between audio and lyrics representation; a certain degree of language-specificity caused by acoustic differences across languages; and the fact that most successful methods in the field are not suited to work in real-time. Real-time lyrics alignment (tracking) would have many useful applications, such as fully automated subtitle display in live concerts and opera. In this work, we describe the first real-time-capable audio-to-lyrics alignment pipeline that is able to robustly track the lyrics of different languages, without additional language information. The proposed model predicts, for each audio frame, a probability vector over (European) phoneme classes, using a very small temporal context, and aligns this vector with a phoneme posteriogram matrix computed beforehand from another recording of the same work, which serves as a reference and a proxy to the written-out lyrics. We evaluate our system's tracking accuracy on the challenging genre of classical opera. Finally, robustness to out-of-training languages is demonstrated in an experiment on Jingju (Beijing opera).

📄 PDF Abstract BibTeX arXiv:2107.14496

Code (0)

등록된 구현이 없습니다.

Tasks

Specificity

Similar Papers 제목 키워드 기반

Unsupervised Generative Adversarial Alignment Representation for Sheet music, Audio and Lyrics

2020-07-29 · Donghuo Zeng, Yi Yu, Keizo Oyama

Sheet music, audio, and lyrics are three main modalities during writing a song. In this paper, we propose an unsupervised generative adversarial alignment representation (UGAAR) model to learn deep discriminative represe…

Representation Learning

DALI: a large Dataset of synchronized Audio, LyrIcs and notes, automatically created using teacher-student machine learning paradigm

2019-06-25 · Gabriel Meseguer-Brocal, Alice Cohen-Hadria, Geoffroy Peeters

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The secon…

A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance

2024-01-17 · Jiyun Park, Sangeon Yong, Taegyun Kwon, Juhan Nam

The goal of real-time lyrics alignment is to take live singing audio as input and to pinpoint the exact position within given lyrics on the fly. The task can benefit real-world applications such as the automatic subtitli…

MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation

2025-05-24 · Woohyun Cho, Youngmin Kim, Sunghyun Lee, Youngjae Yu

Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated musicals, the challenge intensifies due to alignment with visual and audit…

RhythmTranslation

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

2025-07-28 · Renhang Liu, Chia-Yu Hung, Navonil Majumder, Taylor Gautreaux 외 arxiv

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech a…

Audio Generation