paper-with-me

Papers

Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

2024-12-18 · Yunbin Tu, Liang Li, Li Su, Qingming Huang

Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval, and leverages it as a feature extractor for other three challenging tasks solved in a multi-task learning paradigm. Nevertheless, this work struggles to learn the comprehensive cognition of user-preferred content, due to disregarding the hierarchies and association relations across modalities. In this paper, guided by the shallow-to-deep principle, we propose a query-centric audio-visual cognition (QUAG) network to construct a reliable multi-modal representation for moment retrieval, segmentation and step-captioning. Specifically, we first design the modality-synergistic perception to obtain rich audio-visual content, by modeling global contrastive alignment and local fine-grained interaction between visual and audio modalities. Then, we devise the query-centric cognition that uses the deep-level query to perform the temporal-channel filtration on the shallow-level audio-visual representation. This can cognize user-preferred content and thus attain a query-centric audio-visual representation for three tasks. Extensive experiments show QUAG achieves the SOTA results on HIREST. Further, we test QUAG on the query-based video summarization task and verify its good generalization.

📄 PDF Abstract BibTeX arXiv:2412.13543

Code (0)

등록된 구현이 없습니다.

Tasks

Moment RetrievalMulti-Task LearningRetrievalVideo RetrievalVideo Summarization

Similar Papers 제목 키워드 기반

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

2025-01-01 · CVPR 2025 1 · Sihong Huang, Jiaxin Wu, XiaoYong Wei, Yi Cai 외

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-perso…

Action RecognitionScene RecognitionVideo Alignment

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

2025-01-01 · CVPR 2025 1 · Shaofei Huang, Rui Ling, Tianrui Hui, Hongyu Li 외

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where ob…

Learning State-Aware Visual Representations from Audible Interactions

2022-09-27 · Himangi Mittal, Pedro Morgado, Unnat Jain, Abhinav Gupta

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their da…

Action AnticipationAction RecognitionLong Term Action AnticipationObject State Change Classification+1

Character-Centric Understanding of Animated Movies

2025-09-15 · Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie 외 arxiv

Animated movies are captivating for their unique character designs and imaginative storytelling, yet they pose significant challenges for existing recognition systems. Unlike the consistent visual patterns detected by co…

Face Recognition

Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions

2023-08-18 · PLOS ONE 2024 4 · Michael Joannou, Pia Rotshtein, Uta Noppeney

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Tim…