Adversarial Learning for Improved Onsets and Frames Music Transcription
Automatic music transcription is considered to be one of the hardest problems in music information retrieval, yet recent deep learning approaches have achieved substantial improvements on transcription performance. These approaches commonly employ supervised learning models that predict various time-frequency representations, by minimizing element-wise losses such as the cross entropy function. However, applying the loss in this manner assumes conditional independence of each label given the input, and thus cannot accurately express inter-label dependencies. To address this issue, we introduce an adversarial training scheme that operates directly on the time-frequency representations and makes the output distribution closer to the ground-truth. Through adversarial learning, we achieve a consistent improvement in both frame-level and note-level metrics over Onsets and Frames, a state-of-the-art music transcription model. Our results show that adversarial learning can significantly reduce the error rate while increasing the confidence of the model estimations. Our approach is generic and applicable to any transcription model based on multi-label predictions, which are very common in music signal analysis.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMusic Information RetrievalMusic TranscriptionRetrievalSimilar Papers 제목 키워드 기반
Onsets and Frames: Dual-Objective Piano Transcription
We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset eve…
Music TranscriptionTowards Robust Transcription: Exploring Noise Injection Strategies for Training Data Augmentation
Recent advancements in Automatic Piano Transcription (APT) have significantly improved system performance, but the impact of noisy environments on the system performance remains largely unexplored. This study investigate…
Data AugmentationDevelopment of Large Annotated Music Datasets using HMM-based Forced Viterbi Alignment
Datasets are essential for any machine learning task. Automatic Music Transcription (AMT) is one such task, where considerable amount of data is required depending on the way the solution is achieved. Considering the fac…
Music TranscriptionFrom Audio to Symbolic Encoding
Automatic music transcription (AMT) aims to convert raw audio to symbolic music representation. As a fundamental problem of music information retrieval (MIR), AMT is considered a difficult task even for trained human exp…
Information RetrievalMusic Information RetrievalMusic TranscriptionRetrieval+2A Phoneme-Informed Neural Network Model for Note-Level Singing Transcription
Note-level automatic music transcription is one of the most representative music information retrieval (MIR) tasks and has been studied for various instruments to understand music. However, due to the lack of high-qualit…
Information RetrievalMusic Information RetrievalMusic TranscriptionOnset Detection+1