paper-with-me

Papers

MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition

2024-07-08 · Pedro Lima Louro, Hugo Redinho, Ricardo Santos, Ricardo Malheiro, Renato Panda, Rui Pedro Paiva

The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal ensembles that combine audio and lyrics. However, a lack of public, sizable and quality-controlled bimodal databases has hampered the development and improvement of bimodal audio-lyrics systems. This article proposes three new audio, lyrics, and bimodal MER research datasets, collectively referred to as MERGE, which were created using a semi-automatic approach. To comprehensively assess the proposed datasets and establish a baseline for benchmarking, we conducted several experiments for each modality, using feature engineering, machine learning, and deep learning methodologies. Additionally, we propose and validate fixed train-validation-test splits. The obtained results confirm the viability of the proposed datasets, achieving the best overall result of 81.74\% F1-score for bimodal classification.

📄 PDF Abstract BibTeX arXiv:2407.06060

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDeep LearningEmotion RecognitionFeature EngineeringMusic Emotion Recognition

Similar Papers 제목 키워드 기반

Generation of lyrics lines conditioned on music audio clips

2020-09-30 · NLP4MusA 2020 10 · Olga Vechtomova, Gaurav Sahu, Dhruv Kumar

We present a system for generating novel lyrics lines conditioned on music audio. A bimodal neural network model learns to generate lines conditioned on any given short audio clip. The model consists of a spectrogram var…

On-Line Audio-to-Lyrics Alignment Based on a Reference Performance

2021-07-30 · Charles Brazier, Gerhard Widmer

Audio-to-lyrics alignment has become an increasingly active research task in MIR, supported by the emergence of several open-source datasets of audio recordings with word-level lyrics annotations. However, there are stil…

Specificity

MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation

2025-05-24 · Woohyun Cho, Youngmin Kim, Sunghyun Lee, Youngjae Yu

Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated musicals, the challenge intensifies due to alignment with visual and audit…

RhythmTranslation

Exploiting Synchronized Lyrics And Vocal Features For Music Emotion Detection

2019-01-15 · Loreto Parisi, Simone Francia, Silvio Olivastri, Maria Stella Tavella

One of the key points in music recommendation is authoring engaging playlists according to sentiment and emotions. While previous works were mostly based on audio for music discovery and playlists generation, we take adv…

Information RetrievalMusic Information RetrievalMusic RecommendationRetrieval

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion

2025-06-19 · Markus Frohmann, Gabriel Meseguer-Brocal, Markus Schedl, Elena V. Epure

The rapid advancement of AI-based music generation tools is revolutionizing the music industry but also posing challenges to artists, copyright holders, and providers alike. This necessitates reliable methods for detecti…

Music Generation