paper-with-me

Papers

Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision

2018-11-09 · Sanjeel Parekh, Alexey Ozerov, Slim Essid, Ngoc Duong, Patrick Pérez, Gaël Richard

We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object classification in noisy acoustic environments and integrate audio source enhancement capability. This is made possible by a novel use of non-negative matrix factorization for the audio modality. Our approach is founded on the multiple instance learning paradigm. Its effectiveness is established through experiments over a challenging dataset of music instrument performance videos. We also show encouraging visual object localization results.

📄 PDF Abstract BibTeX arXiv:1811.04000

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationMultiple Instance LearningObjectObject LocalizationRepresentation Learning

Similar Papers 제목 키워드 기반

Move2Hear: Active Audio-Visual Source Separation

2021-05-15 · ICCV 2021 10 · Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple …

Audio Source SeparationObject

Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization

2023-08-11 · Sung Jin Um, DongJin Kim, Jung Uk Kim

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, e…

Sound Source Localization

Learning to Separate Object Sounds by Watching Unlabeled Video

2018-04-05 · ECCV 2018 9 · Ruohan Gao, Rogerio Feris, Kristen Grauman

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources toge…

Audio DenoisingAudio Source SeparationDenoisingMulti-Label Learning

Developing an AI-Guided Assistant Device for the Deaf and Hearing Impaired

2025-07-16 · Jiayu, Liu arxiv

This study aims to develop a deep learning system for an accessibility device for the deaf or hearing impaired. The device will accurately localize and identify sound sources in real time. This study will fill an importa…

Audio ClassificationVisual LocalizationObject Detection

A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition

2023-05-30 · Shentong Mo, Pedro Morgado

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independe…

audio-visual learning