Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation
Piano performance is a multimodal activity that intrinsically combines physical actions with the acoustic rendition. Despite growing research interest in analyzing the multimodal nature of piano performance, the laborious process of acquiring large-scale multimodal data remains a significant bottleneck, hindering further progress in this field. To overcome this barrier, we present an integrated web toolkit comprising two graphical user interfaces (GUIs): (i) PiaRec, which supports the synchronized acquisition of audio, video, MIDI, and performance metadata. (ii) ASDF, which enables the efficient annotation of performer fingering from the visual data. Collectively, this system can streamline the acquisition of multimodal piano performance datasets.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
PianoVAM: A Multimodal Piano Performance Dataset
The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano perf…
Information RetrievalHand Pose EstimationPIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
While piano music has become a significant area of study in Music Information Retrieval (MIR), there is a notable lack of datasets for piano solo music with text labels. To address this gap, we present PIAST (PIano datas…
Information RetrievalMusic Information RetrievalMusic TaggingRetrieval+1Piano Skills Assessment
Can a computer determine a piano player's skill level? Is it preferable to base this assessment on visual analysis of the player's performance or should we trust our ears over our eyes? Since current CNNs have difficulty…
Action Quality AssessmentAudio ClassificationMultimodal Deep LearningSkills Assessment+3Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, existing evaluation datasets do not fully…
Music GenerationLecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
Evaluating the quality of slide-based multimedia instruction is challenging. Existing methods like manual assessment, reference-based metrics, and large language model evaluators face limitations in scalability, context …
Language ModelingLanguage ModellingLarge Language Model