paper-with-me

Papers

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

2025-12-23 · Jingqi Tian, Yiheng Du, Haoji Zhang, Yuji Wang, Isaac Ning Lee, Xulong Bai, Tianrui Zhu, Jingxuan Niu, Yansong Tang arxiv

Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existing methods often struggle with multi-source entanglement and audio-visual misalignment, leading to a dominance bias toward acoustically or visually salient objects (i.e., louder or larger ones) at the expense of subtler or co-occurring sources. To address these challenges, we propose DDAVS: Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation. To mitigate multi-source entanglement, DDAVS employs learnable queries to extract audio semantics and anchor them within a structured semantic space derived from an audio prototype memory bank. This process is further optimized through contrastive learning to enhance discriminability and robustness. To alleviate audio-visual misalignment, DDAVS introduces dual cross attention with delayed modality interaction, improving the robustness of multimodal alignment. Extensive experiments on the AVS-Objects and VPO benchmarks demonstrate that DDAVS achieves state-of-the-art performance across single-source, multi-source, and multi-class multi-instance scenarios. These results validate the effectiveness and generalization ability of our framework under challenging real-world audio-visual segmentation conditions. Project page: https://trilarflagz.github.io/DDAVS-page/

📄 PDF Abstract BibTeX arXiv:2512.20117

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Talking Face Generation by Adversarially Disentangled Audio-Visual Representation

2018-07-20 · Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo 외

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the s…

Face GenerationLip ReadingRetrievalTalking Face Generation+1

Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network

2019-08-30 · ICML 2020 1 · Javier S. Turek, Shailee Jain, Vy Vo, Mihai Capota 외

Recent work has shown that topological enhancements to recurrent neural networks (RNNs) can increase their expressiveness and representational capacity. Two popular enhancements are stacked RNNs, which increases the capa…

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

2025-09-30 · Chetwin Low, Weimin Wang, Calder Katyal arxiv

Audio-video generation has often relied on complex multi-stage architectures or sequential synthesis of sound and visuals. We introduce Ovi, a unified paradigm for audio-video generation that models the two modalities as…

Video Generation

Boosting Latent Diffusion Models via Disentangled Representation Alignment

2026-01-09 · John Page, Xuesong Niu, Kai Wu, Kun Gai arxiv

Latent Diffusion Models (LDMs) rely heavily on the compressed latent space provided by Variational Autoencoders (VAEs) for high-quality image generation. Recent studies have attempted to obtain generation-friendly VAEs b…

Image Generation

QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition

2023-09-29 · CVPR 2024 1 · Xiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng 외

Audiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved, establishi…

Quantization