paper-with-me

Papers

V-SlowFast Network for Efficient Visual Sound Separation

2021-09-18 · Lingyu Zhu, Esa Rahtu

The objective of this paper is to perform visual sound separation: i) we study visual sound separation on spectrograms of different temporal resolutions; ii) we propose a new light yet efficient three-stream framework V-SlowFast that operates on Visual frame, Slow spectrogram, and Fast spectrogram. The Slow spectrogram captures the coarse temporal resolution while the Fast spectrogram contains the fine-grained temporal resolution; iii) we introduce two contrastive objectives to encourage the network to learn discriminative visual features for separating sounds; iv) we propose an audio-visual global attention module for audio and visual feature fusion; v) the introduced V-SlowFast model outperforms previous state-of-the-art in single-frame based visual sound separation on small- and large-scale datasets: MUSIC-21, AVE, and VGG-Sound. We also propose a small V-SlowFast architecture variant, which achieves 74.2% reduction in the number of model parameters and 81.4% reduction in GMACs compared to the previous multi-stage models. Project page: https://ly-zhu.github.io/V-SlowFast

📄 PDF Abstract BibTeX arXiv:2109.08867

Code (1)

ly-zhu/ly-zhu.github.io

Similar Papers 제목 키워드 기반

Audiovisual SlowFast Networks for Video Recognition

2020-01-23 · Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik 외

We present Audiovisual SlowFast Networks, an architecture for integrated audiovisual perception. AVSlowFast has Slow and Fast visual pathways that are deeply integrated with a Faster Audio pathway to model vision and sou…

Action ClassificationVideo Recognition

MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization

2025-09-24 · Jianxuan Yang, Xiaoran Yang, Lipan Zhang, Xinyue Guo 외 arxiv

Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical limitations. First, existing methods fac…

Audio Generation

High-Quality Visually-Guided Sound Separation from Diverse Categories

2023-07-31 · Chao Huang, Susan Liang, Yapeng Tian, Anurag Kumar 외

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-bas…

Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation

2023-10-18 · Yiyang Su, Ali Vosoughi, Shijian Deng, Yapeng Tian 외

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduce…

cross-modal alignment

Leveraging Category Information for Single-Frame Visual Sound Source Separation

2020-07-15 · Lingyu Zhu, Esa Rahtu

Visual sound source separation aims at identifying sound components from a given sound mixture with the presence of visual cues. Prior works have demonstrated impressive results, but with the expense of large multi-stage…

Optical Flow Estimation