Data Cleansing with Contrastive Learning for Vocal Note Event Annotations
Data cleansing is a well studied strategy for cleaning erroneous labels in datasets, which has not yet been widely adopted in Music Information Retrieval. Previously proposed data cleansing models do not consider structured (e.g. time varying) labels, such as those common to music data. We propose a novel data cleansing model for time-varying, structured labels which exploits the local structure of the labels, and demonstrate its usefulness for vocal note event annotations in music. %Our model is trained in a contrastive learning manner by automatically creating local deformations of likely correct labels. Our model is trained in a contrastive learning manner by automatically contrasting likely correct labels pairs against local deformations of them. We demonstrate that the accuracy of a transcription model improves greatly when trained using our proposed strategy compared with the accuracy when trained using the original dataset. Additionally we use our model to estimate the annotation error rates in the DALI dataset, and highlight other potential uses for this type of model.
Code (1)
Tasks
Contrastive LearningInformation RetrievalMusic Information RetrievalRetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Metrical-accent Aware Vocal Onset Detection in Polyphonic Audio
The goal of this study is the automatic detection of onsets of the singing voice in polyphonic audio recordings. Starting with a hypothesis that the knowledge of the current position in a metrical cycle (i.e. metrical ac…
Onset DetectionPositionAutomatic recognition of element classes and boundaries in the birdsong with variable sequences
Researches on sequential vocalization often require analysis of vocalizations in long continuous sounds. In such studies as developmental ones or studies across generations in which days or months of vocalizations must b…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Boundary DetectionGeneral Classification+2Phoneme-Informed Note Segmentation of Monophonic Vocal Music
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sound concepts through voice, QBV offers th…
Contrastive LearningAutomatic Transcription of Flamenco Singing from Polyphonic Music Recordings
Automatic note-level transcription is considered one of the most challenging tasks in music information retrieval. The specific case of flamenco singing transcription poses a particular challenge due to its complex melod…
Information RetrievalMusic Information RetrievalOnset DetectionRetrieval