paper-with-me

홈 › Papers

Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual Segmentation

2024-12-11 · Kexin Li, Zongxin Yang, Yi Yang, Jun Xiao

Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects that accurately align with the corresponding audio. However, existing methods often face temporal misalignment, where audio cues and segmentation results are not temporally coordinated. Audio provides two critical pieces of information: i) target object-level details and ii) the timing of when objects start and stop producing sounds. Current methods focus more on object-level information but neglect the boundaries of audio semantic changes, leading to temporal misalignment. To address this issue, we propose a Collaborative Hybrid Propagator Framework~(Co-Prop). This framework includes two main steps: Preliminary Audio Boundary Anchoring and Frame-by-Frame Audio-Insert Propagation. To Anchor the audio boundary, we employ retrieval-assist prompts with Qwen large language models to identify control points of audio semantic changes. These control points split the audio into semantically consistent audio portions. After obtaining the control point lists, we propose the Audio Insertion Propagator to process each audio portion using a frame-by-frame audio insertion propagation and matching approach. We curated a compact dataset comprising diverse source conversion cases and devised a metric to assess alignment rates. Compared to traditional simultaneous processing methods, our approach reduces memory requirements and facilitates frame alignment. Experimental results demonstrate the effectiveness of our approach across three datasets and two backbones. Furthermore, our method can be integrated with existing AVVS approaches, offering plug-and-play functionality to enhance their performance.

📄 PDF Abstract BibTeX arXiv:2412.08161

Code (0)

등록된 구현이 없습니다.

Tasks

Video SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Focus 설명 없음

Similar Papers 제목 키워드 기반

CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment

2026-02-09 · Yunzuo Hu, Wen Li, Jing Zhang arxiv

Audio-visual learning suffers from modality misalignment caused by off-screen sources and background clutter, and current methods usually amplify irrelevant regions or moments, leading to unstable training and degraded r…

Long-Video Audio Synthesis with Multi-Agent Collaboration

2025-03-13 · Yehang Zhang, Xinli Xu, Xiaojie Xu, Li Liu 외

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form c…

Audio SynthesisScene SegmentationScript Generation

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

2026-08-03 · Jiayu Chen, Xiaoyu Wu, Rongshan Gao, Maoliang Li 외 arxiv

Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due to the iterative denoising process of dif…

Video Generation

MAVIN: Multi-Shot Audio-Visual Generation with Narrative Control

2026-06-28 · Kaiqi Liu, Yunyao Mao, Ziqi Cai, Zheng Geng 외 arxiv

While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing methods suffer from temporal misalignment, …

CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors

2026-08-21 · Chi Li, Rui Lin, Aobo Ji, Dongzhu Xu arxiv

Collaborative perception extends the sensing range of a single vehicle by fusing observations from nearby agents, which improves the robustness of autonomous driving. In realistic deployments, however, the received colla…

Autonomous Driving