paper-with-me

홈 › Papers

HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training

2022-12-30 · ICCV 2023 1 · Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, Fei Huang

Video-language pre-training has advanced the performance of various downstream video-language tasks. However, most previous methods directly inherit or adapt typical image-language pre-training paradigms to video-language pre-training, thus not fully exploiting the unique characteristic of video, i.e., temporal. In this paper, we propose a Hierarchical Temporal-Aware video-language pre-training framework, HiTeA, with two novel pre-training tasks for modeling cross-modal alignment between moments and texts as well as the temporal relations of video-text pairs. Specifically, we propose a cross-modal moment exploration task to explore moments in videos, which results in detailed video moment representation. Besides, the inherent temporal relations are captured by aligning video-text pairs as a whole in different time resolutions with multi-modal temporal relation exploration task. Furthermore, we introduce the shuffling test to evaluate the temporal reliance of datasets and video-language pre-training models. We achieve state-of-the-art results on 15 well-established video-language understanding and generation tasks, especially on temporal-oriented datasets (e.g., SSv2-Template and SSv2-Label) with 8.6% and 11.1% improvement respectively. HiTeA also demonstrates strong generalization ability when directly transferred to downstream tasks in a zero-shot manner. Models and demo will be available on ModelScope.

📄 PDF Abstract BibTeX arXiv:2212.14546

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentTGIF-ActionTGIF-FrameTGIF-TransitionVideo CaptioningVideo Question AnsweringVideo RetrievalVisual Question Answering (VQA)Zero-Shot LearningZero-Shot Video Retrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Object-aware Aggregation with Bidirectional Temporal Graph for Video Captioning

2019-06-11 · CVPR 2019 6 · Junchao Zhang, Yuxin Peng

Video captioning aims to automatically generate natural language descriptions of video content, which has drawn a lot of attention recent years. Generating accurate and fine-grained captions needs to not only understand …

ObjectVideo Captioning

VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree

2025-10-26 · Wenlong Li, Yifei Xu, Yuan Rao, Zhenhua Wang 외 arxiv

Video anomaly detection (VAD) focuses on identifying anomalies in videos. Supervised methods demand substantial in-domain training data and fail to deliver clear explanations for anomalies. In contrast, training-free met…

Generic Event Boundary DetectionVideo Anomaly Detection

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

2026-06-19 · Awais Rauf, Ahmed Hasssan, Greg Slabaugh arxiv

Understanding long videos requires fine-grained perception and multi-step, higher-order reasoning over complex, long-range spatio-temporal dynamics. Vision-language models (VLMs) encode video frames into visual tokens an…

Relational ReasoningSemantic Retrieval

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video

2026-04-13 · Junfu Pu, Yuxin Chen, Teng Wang, Ying Shan arxiv

Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detailed, temporally grounded scripts remains …

Reinforcement Learning

VidLA: Video-Language Alignment at Scale

2024-03-21 · CVPR 2024 1 · Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan, Son Tran 외

In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do not capture both short-range and long-ra…

Language ModellingVisual Grounding