The Sound of Bounding-Boxes
In the task of audio-visual sound source separation, which leverages visual information for sound source separation, identifying objects in an image is a crucial step prior to separating the sound source. However, existing methods that assign sound on detected bounding boxes suffer from a problem that their approach heavily relies on pre-trained object detectors. Specifically, when using these existing methods, it is required to predetermine all the possible categories of objects that can produce sound and use an object detector applicable to all such categories. To tackle this problem, we propose a fully unsupervised method that learns to detect objects in an image and separate sound source simultaneously. As our method does not rely on any pre-trained detector, our method is applicable to arbitrary categories without any additional annotation. Furthermore, although being fully unsupervised, we found that our method performs comparably in separation accuracy.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes
To tackle sound event detection (SED), we propose frequency dependent networks (FreDNets), which heavily leverage frequency-dependent methods. We apply frequency warping and FilterAugment, which are frequency-dependent d…
Change DetectionData AugmentationEvent DetectionPseudo Label+1Sound Event Bounding Boxes
Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence confidence in short time frames. Then, thre…
Change DetectionEvent DetectionPredictionSound Event DetectionMedical image segmentation with imperfect 3D bounding boxes
The development of high quality medical image segmentation algorithms depends on the availability of large datasets with pixel-level labels. The challenges of collecting such datasets, especially in case of 3D volumes, m…
Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation+1False Negative/Positive Control for SAM on Noisy Medical Images
The Segment Anything Model (SAM) is a recently developed all-range foundation model for image segmentation. It can use sparse manual prompts such as bounding boxes to generate pixel-level segmentation in natural images b…
Image SegmentationMedical Image SegmentationSegmentationSemantic SegmentationSemantic Segmentation and Object Detection Towards Instance Segmentation: Breast Tumor Identification
Breast cancer is one of the factors that cause the increase of mortality of women. The most widely used method for diagnosing this geological disease i.e. breast cancer is the ultrasound scan. Several key features such a…
DecoderInstance Segmentationobject-detectionObject Detection+2