Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
This study introduces a meta-learning-based approach for low-resource Tabla Stroke Transcription (TST) and $t\bar{a}la$ identification in Hindustani classical music. Using Model-Agnostic Meta-Learning (MAML), we address the challenges of limited annotated datasets and label heterogeneity, enabling rapid adaptation to new tasks with minimal data. The method is validated across various datasets, including tabla solo and concert recordings, demonstrating robustness in polyphonic audio scenarios. We propose two novel $t\bar{a}la$ identification techniques based on stroke sequences and rhythmic patterns. Additionally, the approach proves effective for Automatic Drum Transcription (ADT), showcasing its flexibility for Indian and Western percussion music. Experimental results show that the proposed method outperforms existing techniques in low-resource settings, significantly contributing to music transcription and studying musical traditions through computational tools.
Code (0)
등록된 구현이 없습니다.
Tasks
Drum TranscriptionMeta-LearningMusic TranscriptionSimilar Papers 제목 키워드 기반
Automatic Transcription of Drum Strokes in Carnatic Music
The mridangam is a double-headed percussion instrument that plays a key role in Carnatic music concerts. This paper presents a novel automatic transcription algorithm to classify the strokes played on the mridangam. Onse…
Onset DetectionVulnerability in Acquisition, Language Impairments in Dutch: Creating a VALID Data Archive
The VALID Data Archive is an open multimedia data archive (under construction) with data from speakers suffering from language impairments. We report on a pilot project in the CLARIN-NL framework in which five data resou…
AllLanguage AcquisitionvalidArabic Dialect Identification Using iVectors and ASR Transcripts
This paper presents the systems submitted by the MAZA team to the Arabic Dialect Identification (ADI) shared task at the VarDial Evaluation Campaign 2017. The goal of the task is to evaluate computational models to ident…
Dialect IdentificationMachine TranslationScalable Music Cover Retrieval Using Lyrics-Aligned Audio Embeddings
Music Cover Retrieval, also known as Version Identification, aims to recognize distinct renditions of the same underlying musical work, a task central to catalog management, copyright enforcement, and music retrieval. St…
Computational EfficiencyDeep Embeddings for Robust User-Based Amateur Vocal Percussion Classification
Vocal Percussion Transcription (VPT) is concerned with the automatic detection and classification of vocal percussion sound events, allowing music creators and producers to sketch drum lines on the fly. Classifier algori…
Classificationfeature selectionspeech-recognitionSpeech Recognition