paper-with-me

홈 › Papers

Music Gesture for Visual Sound Separation

2020-04-20 · CVPR 2020 6 · Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, Antonio Torralba

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited abilities to find the correlations between audio signals and visual points, especially when separating multiple instruments of the same types, such as multiple violins in a scene. To address this, we propose "Music Gesture," a keypoint-based structured representation to explicitly model the body and finger movements of musicians when they perform music. We first adopt a context-aware graph network to integrate visual semantic context with body dynamics, and then apply an audio-visual fusion model to associate body movements with the corresponding audio signals. Experimental results on three music performance datasets show: 1) strong improvements upon benchmark metrics for hetero-musical separation tasks (i.e. different instruments); 2) new ability for effective homo-musical separation for piano, flute, and trumpet duets, which to our best knowledge has never been achieved with alternative methods. Project page: http://music-gesture.csail.mit.edu.

📄 PDF Abstract BibTeX arXiv:2004.09476

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow Estimation

Similar Papers 제목 키워드 기반

Visual Scene Graphs for Audio Source Separation

2021-09-24 · ICCV 2021 10 · Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop Cherian

State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sou…

Audio Source SeparationVisually Guided Sound Source Separation

Mugeetion: Musical Interface Using Facial Gesture and Emotion

2018-09-14 · Eunjeong Stella Koh, Shahrokh Yadegari

People feel emotions when listening to music. However, emotions are not tangible objects that can be exploited in the music composition process as they are difficult to capture and quantify in algorithms. We present a no…

High-Quality Visually-Guided Sound Separation from Diverse Categories

2023-07-31 · Chao Huang, Susan Liang, Yapeng Tian, Anurag Kumar 외

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-bas…

Semantic Grouping Network for Audio Source Separation

2024-07-04 · Shentong Mo, Yapeng Tian

Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extracted high-level semantics from visual in…

Audio Source Separation

SAM Audio: Segment Anything in Audio

2025-12-19 · Bowen Shi, Andros Tjandra, John Hoffman, Helin Wang 외 arxiv

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific,…

Audio Source Separation