paper-with-me

Papers

Visually Guided Sound Source Separation using Cascaded Opponent Filter Network

2020-06-04 · Lingyu Zhu, Esa Rahtu

The objective of this paper is to recover the original component signals from a mixture audio with the aid of visual cues of the sound sources. Such task is usually referred as visually guided sound source separation. The proposed Cascaded Opponent Filter (COF) framework consists of multiple stages, which recursively refine the source separation. A key element in COF is a novel opponent filter module that identifies and relocates residual components between sources. The system is guided by the appearance and motion of the source, and, for this purpose, we study different representations based on video frames, optical flows, dynamic images, and their combinations. Finally, we propose a Sound Source Location Masking (SSLM) technique, which, together with COF, produces a pixel level mask of the source location. The entire system is trained end-to-end using a large set of unlabelled videos. We compare COF with recent baselines and obtain the state-of-the-art performance in three challenging datasets (MUSIC, A-MUSIC, and A-NATURAL). Project page: https://ly-zhu.github.io/cof-net.

📄 PDF Abstract BibTeX arXiv:2006.03028

Code (1)

ly-zhu/ly-zhu.github.io

Tasks

Visually Guided Sound Source Separation

Similar Papers 제목 키워드 기반

Weakly-supervised Audio-visual Sound Source Detection and Separation

2021-03-25 · Tanzila Rahman, Leonid Sigal

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mi…

Audio Source SeparationDenoisingObjectSegmentation+3

Co-Separating Sounds of Visual Objects

2019-04-16 · ICCV 2019 10 · Ruohan Gao, Kristen Grauman

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificial…

Audio DenoisingAudio Source SeparationDenoising

High-Quality Visually-Guided Sound Separation from Diverse Categories

2023-07-31 · Chao Huang, Susan Liang, Yapeng Tian, Anurag Kumar 외

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-bas…

Visually-Guided Sound Source Separation with Audio-Visual Predictive Coding

2023-06-19 · Zengjie Song, Zhaoxiang Zhang

The framework of visually-guided sound source separation generally consists of three parts: visual feature extraction, multimodal feature fusion, and sound signal processing. An ongoing trend in this field has been to ta…

validVisually Guided Sound Source Separation

Visual Scene Graphs for Audio Source Separation

2021-09-24 · ICCV 2021 10 · Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop Cherian

State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sou…

Audio Source SeparationVisually Guided Sound Source Separation