paper-with-me

홈 › Papers

Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition

2024-01-09 · Xuzheng Yu, Chen Jiang, Wei zhang, Tian Gan, Linlin Chao, Jianan Zhao, Yuan Cheng, Qingpei Guo, Wei Chu

With the explosive growth of video data in real-world applications, a comprehensive representation of videos becomes increasingly important. In this paper, we address the problem of video scene recognition, whose goal is to learn a high-level video representation to classify scenes in videos. Due to the diversity and complexity of video contents in realistic scenarios, this task remains a challenge. Most existing works identify scenes for videos only from visual or textual information in a temporal perspective, ignoring the valuable information hidden in single frames, while several earlier studies only recognize scenes for separate images in a non-temporal perspective. We argue that these two perspectives are both meaningful for this task and complementary to each other, meanwhile, externally introduced knowledge can also promote the comprehension of videos. We propose a novel two-stream framework to model video representations from multiple perspectives, i.e. temporal and non-temporal perspectives, and integrate the two perspectives in an end-to-end manner by self-distillation. Besides, we design a knowledge-enhanced feature fusion and label prediction method that contributes to naturally introducing knowledge into the task of video scene recognition. Experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method.

📄 PDF Abstract BibTeX arXiv:2401.04354

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningScene Recognition

Similar Papers 제목 키워드 기반

4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models

2024-06-11 · Heng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace 외

Existing dynamic scene generation methods mostly rely on distilling knowledge from pre-trained 3D generative models, which are typically fine-tuned on synthetic object datasets. As a result, the generated scenes are ofte…

Scene GenerationVideo Generation

Light-VQA: A Multi-Dimensional Quality Assessment Model for Low-Light Video Enhancement

2023-05-16 · Yunlong Dong, Xiaohong Liu, Yixuan Gao, Xunchu Zhou 외

Recently, Users Generated Content (UGC) videos becomes ubiquitous in our daily lives. However, due to the limitations of photographic equipments and techniques, UGC videos often contain various degradations, in which one…

Video EnhancementVideo Quality AssessmentVisual Question Answering (VQA)

CLOP: Video-and-Language Pre-Training with Knowledge Regularizations

2022-11-07 · Guohao Li, Hu Yang, Feng He, Zhifan Feng 외

Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner, without considering explicit structural…

Contrastive LearningRetrievalVideo Retrieval

Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval

2026-03-25 · Junkai Yang, Qirui Wang, Yaoqing Jin, Shuai Ma 외 arxiv

Retrieving partially relevant segments from untrimmed videos remains difficult due to two persistent challenges: the mismatch in information density between text and video segments, and limited attention mechanisms that …

Partially Relevant Video Retrieval

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

2026-08-19 · Mohammad Zamani, Fatemeh Ziaeetabar arxiv

Egocentric video captures activities from the wearer's perspective, providing a direct view of human attention, hand--object interaction, and goal-directed behavior. This perspective is increasingly important for wearabl…

Representation LearningDomain GeneralizationDecision Making