Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision
Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on multimodal unsupervised learning (as infants) and active learning (as children). With this motivation, we present a learning framework for sound representation and recognition that combines (i) a self-supervised objective based on a general notion of unimodal and cross-modal coincidence, (ii) a clustering objective that reflects our need to impose categorical structure on our experiences, and (iii) a cluster-based active learning procedure that solicits targeted weak supervision to consolidate categories into relevant semantic classes. By training a combined sound embedding/clustering/classification network according to these criteria, we achieve a new state-of-the-art unsupervised audio representation and demonstrate up to a 20-fold reduction in the number of labels required to reach a desired classification performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningClusteringGeneral ClassificationSimilar Papers 제목 키워드 기반
Neural coincidence detection strategies during perception of multi-pitch musical tones
Multi-pitch perception is investigated in a listening test using 30 recordings of musical sounds with two tones played simultaneously, except for two gong sounds with inharmonic overtone spectrum, judging roughness and s…
Fuzzy Core Equivalence in Large Economies: A Role for the Infinite-Dimensional Lyapunov Theorem
We present the equivalence between the fuzzy core and the core under minimal assumptions. Due to the exact version of the Lyapunov convexity theorem in Banach spaces, we clarify that the additional structure of commodity…
Unsupervised learning human's activities by overexpressed recognized non-speech sounds
Human activity and environment produces sounds such as, at home, the noise produced by water, cough, or television. These sounds can be used to determine the activity in the environment. The objective is to monitor a per…
TAGMinimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich Networks
Text categorization is an essential task in Web content analysis. Considering the ever-evolving Web data and new emerging categories, instead of the laborious supervised setting, in this paper, we focus on the minimally-…
Product CategorizationText CategorizationInto the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds
Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work,…
Scene Understanding