Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision
We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object classification in noisy acoustic environments and integrate audio source enhancement capability. This is made possible by a novel use of non-negative matrix factorization for the audio modality. Our approach is founded on the multiple instance learning paradigm. Its effectiveness is established through experiments over a challenging dataset of music instrument performance videos. We also show encouraging visual object localization results.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationMultiple Instance LearningObjectObject LocalizationRepresentation LearningSimilar Papers 제목 키워드 기반
Move2Hear: Active Audio-Visual Source Separation
We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple …
Audio Source SeparationObjectAudio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization
The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, e…
Sound Source LocalizationLearning to Separate Object Sounds by Watching Unlabeled Video
Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources toge…
Audio DenoisingAudio Source SeparationDenoisingMulti-Label LearningDeveloping an AI-Guided Assistant Device for the Deaf and Hearing Impaired
This study aims to develop a deep learning system for an accessibility device for the deaf or hearing impaired. The device will accurately localize and identify sound sources in real time. This study will fill an importa…
Audio ClassificationVisual LocalizationObject DetectionA Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition
The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independe…
audio-visual learning