paper-with-me

Papers

Query-aware Long Video Localization and Relation Discrimination for Deep Video Understanding

2023-10-19 · Yuanxing Xu, Yuting Wei, Bin Wu

The surge in video and social media content underscores the need for a deeper understanding of multimedia data. Most of the existing mature video understanding techniques perform well with short formats and content that requires only shallow understanding, but do not perform well with long format videos that require deep understanding and reasoning. Deep Video Understanding (DVU) Challenge aims to push the boundaries of multimodal extraction, fusion, and analytics to address the problem of holistically analyzing long videos and extract useful knowledge to solve different types of queries. This paper introduces a query-aware method for long video localization and relation discrimination, leveraging an imagelanguage pretrained model. This model adeptly selects frames pertinent to queries, obviating the need for a complete movie-level knowledge graph. Our approach achieved first and fourth positions for two groups of movie-level queries. Sufficient experiments and final rankings demonstrate its effectiveness and robustness.

📄 PDF Abstract BibTeX arXiv:2310.12724

Code (0)

등록된 구현이 없습니다.

Tasks

RelationVideo Understanding

Similar Papers 제목 키워드 기반

Single-Stage Visual Query Localization in Egocentric Videos

2023-06-15 · NeurIPS 2023 11

Visual Query Localization on long-form egocentric videos requires spatio-temporal search and localization of visually specified objects and is vital to build episodic memory systems. Prior work develops complex multi-sta…

object-detectionObject DetectionTemporal LocalizationVideo Relationship

DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video

2020-10-13 · Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Basura Fernando, Hongdong Li 외

This paper studies the task of temporal moment localization in a long untrimmed video using natural language query. Given a query sentence, the goal is to determine the start and end of the relevant segment within the vi…

Sentence

Graph Neural Network for Video Relocalization

2020-07-20 · Yuan Zhou, Mingfei Wang, Ruolin Wang, Shuwei Huo

In this paper, we focus on video relocalization task, which uses a query video clip as input to retrieve a semantic relative video clip in another untrimmed long video. we find that in video relocalization datasets, ther…

Graph Neural NetworkMoment Retrieval

Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos

2021-10-12 · Zongmeng Zhang, Xianjing Han, Xuemeng Song, Yan Yan 외

This paper focuses on tackling the problem of temporal language localization in videos, which aims to identify the start and end points of a moment described by a natural language sentence in an untrimmed video. However,…

Semantic correspondenceSemantic SimilaritySemantic Textual SimilaritySentence

Jointly Visual- and Semantic-Aware Graph Memory Networks for Temporal Sentence Localization in Videos

2023-03-02 · Daizong Liu, Pan Zhou

Temporal sentence localization in videos (TSLV) aims to retrieve the most interested segment in an untrimmed video according to a given sentence query. However, almost of existing TSLV approaches suffer from the same lim…

Representation LearningSentenceVisual Reasoning