paper-with-me

Papers

Video Object Segmentation-Aware Audio Generation

2025-09-30 · Ilpo Viertola, Vladimir Iashin, Esa Rahtu arxiv

Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide precise methods for prioritizing a specific object within a scene, generating unnecessary background sounds, or focusing on the wrong objects. To address this gap, we introduce the novel task of video object segmentation-aware audio generation, which explicitly conditions sound synthesis on object-level segmentation maps. We present SAGANet, a new multimodal generative model that enables controllable audio generation by leveraging visual segmentation masks along with video and textual cues. Our model provides users with fine-grained and visually localized control over audio generation. To support this task and further research on segmentation-aware Foley, we propose Segmented Music Solos, a benchmark dataset of musical instrument performance videos with segmentation information. Our method demonstrates substantial improvements over current state-of-the-art methods and sets a new standard for controllable, high-fidelity Foley synthesis. Code, samples, and Segmented Music Solos are available at https://saganet.notion.site

📄 PDF Abstract BibTeX arXiv:2509.26604

Code (0)

등록된 구현이 없습니다.

Tasks

Video Object SegmentationAudio Generation

Similar Papers 제목 키워드 기반

PAVAS: Physics-Aware Video-to-Audio Synthesis

2025-12-09 · Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka, Tae-Hyun Oh 외 arxiv

Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing visual-acoustic correlations without consid…

3D Reconstruction

Audio-Visual Segmentation by Exploring Cross-Modal Mutual Semantics

2023-07-31 · Chen Liu, Peike Li, Xingqun Qi, Hu Zhang 외

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we…

ObjectSegmentationSemantic Segmentation

3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio

2026-03-24 · Jihwan Hong, Jaeyoung Do arxiv

Audio-based Referring Video Object Segmentation (ARVOS) requires grounding audio queries into pixel-level object masks over time, posing challenges in bridging acoustic signals with spatio-temporal visual representations…

Referring Video Object SegmentationVideo Segmentation

Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

2023-07-25 · Jinxiang Liu, Chen Ju, Chaofan Ma, Yanfeng Wang 외

The goal of the audio-visual segmentation (AVS) task is to segment the sounding objects in the video frames using audio cues. However, current fusion-based methods have the performance limitations due to the small recept…

DecoderSegmentation

Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation

2025-09-26 · Jinbae Seo, Hyeongjun Kwon, Kwonyoung Kim, Jiyoung Lee 외 arxiv

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform add…

Instance Segmentation