paper-with-me

Papers

STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment

2023-10-12 · Jaewoo Lee, Jaehong Yoon, Wonjae Kim, Yunji Kim, Sung Ju Hwang

Continuously learning a variety of audio-video semantics over time is crucial for audio-related reasoning tasks in our ever-evolving world. However, this is a nontrivial problem and poses two critical challenges: sparse spatio-temporal correlation between audio-video pairs and multimodal correlation overwriting that forgets audio-video relations. To tackle this problem, we propose a new continual audio-video pre-training method with two novel ideas: (1) Localized Patch Importance Scoring: we introduce a multimodal encoder to determine the importance score for each patch, emphasizing semantically intertwined audio-video patches. (2) Replay-guided Correlation Assessment: to reduce the corruption of previously learned audiovisual knowledge due to drift, we propose to assess the correlation of the current patches on the past steps to identify the patches exhibiting high correlations with the past steps. Based on the results from the two ideas, we perform probabilistic patch selection for effective continual audio-video pre-training. Experimental validation on multiple benchmarks shows that our method achieves a 3.69%p of relative performance gain in zero-shot retrieval tasks compared to strong continual learning baselines, while reducing memory consumption by ~45%.

📄 PDF Abstract BibTeX arXiv:2310.08204

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningRepresentation LearningVideo Alignment

Methods 이 논문이 사용한 방법론

FLAVA FLAVA aims at building a single holistic universal model that targets all modalities at once. FLAVA is a language vision alignment model that learns strong representations from…

Similar Papers 제목 키워드 기반

CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

2025-11-19 · Hokuto Munakata, Takehiro Imamura, Taichi Nishimura, Tatsuya Komatsu arxiv

We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data.…

Moment Retrieval

Continually Evolving Skill Knowledge in Vision Language Action Model

2025-11-22 · Yuxuan Wu, Guangming Wang, Zhiheng Yang, Tianchen Deng 외 arxiv

Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains challenging, especially for efficient adaptation. Existing continual imitation lea…

Continual Learning

Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation

2026-03-09 · Siddeshwar Raghavan, Gautham Vinod, Bruce Coburn, Fengqing Zhu arxiv

Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing …

Continual Learning

R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios

2025-11-21 · Lu Zhu, Tiantian Geng, Yangye Chen, Teng Wang 외 arxiv

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video scenarios, failing to reflect the comple…

Reinforcement LearningVisual Reasoning

ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation

2025-11-15 · Jiahui Sun, Weining Wang, Mingzhen Sun, Yirong Yang 외 arxiv

Sounding Video Generation (SVG) remains a challenging task due to the inherent structural misalignment between audio and video, as well as the high computational cost of multimodal data processing. In this paper, we intr…

Computational EfficiencyVideo Generation