MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition
The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal ensembles that combine audio and lyrics. However, a lack of public, sizable and quality-controlled bimodal databases has hampered the development and improvement of bimodal audio-lyrics systems. This article proposes three new audio, lyrics, and bimodal MER research datasets, collectively referred to as MERGE, which were created using a semi-automatic approach. To comprehensively assess the proposed datasets and establish a baseline for benchmarking, we conducted several experiments for each modality, using feature engineering, machine learning, and deep learning methodologies. Additionally, we propose and validate fixed train-validation-test splits. The obtained results confirm the viability of the proposed datasets, achieving the best overall result of 81.74\% F1-score for bimodal classification.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingDeep LearningEmotion RecognitionFeature EngineeringMusic Emotion RecognitionSimilar Papers 제목 키워드 기반
Generation of lyrics lines conditioned on music audio clips
We present a system for generating novel lyrics lines conditioned on music audio. A bimodal neural network model learns to generate lines conditioned on any given short audio clip. The model consists of a spectrogram var…
On-Line Audio-to-Lyrics Alignment Based on a Reference Performance
Audio-to-lyrics alignment has become an increasingly active research task in MIR, supported by the emergence of several open-source datasets of audio recordings with word-level lyrics annotations. However, there are stil…
SpecificityMAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation
Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated musicals, the challenge intensifies due to alignment with visual and audit…
RhythmTranslationExploiting Synchronized Lyrics And Vocal Features For Music Emotion Detection
One of the key points in music recommendation is authoring engaging playlists according to sentiment and emotions. While previous works were mostly based on audio for music discovery and playlists generation, we take adv…
Information RetrievalMusic Information RetrievalMusic RecommendationRetrievalDouble Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
The rapid advancement of AI-based music generation tools is revolutionizing the music industry but also posing challenges to artists, copyright holders, and providers alike. This necessitates reliable methods for detecti…
Music Generation