paper-with-me

Papers

Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding

2020-08-16 · Zhu Zhang, Zhou Zhao, Zhijie Lin, Baoxing Huai, Nicholas Jing Yuan

Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pairs. In this paper, we explore spatio-temporal video grounding on unaligned data and multi-form sentences. This challenging task requires to capture critical object relations to identify the queried target. However, existing approaches cannot distinguish notable objects and remain in ineffective relation modeling between unnecessary objects. Thus, we propose a novel object-aware multi-branch relation network for object-aware relation discovery. Concretely, we first devise multiple branches to develop object-aware region modeling, where each branch focuses on a crucial object mentioned in the sentence. We then propose multi-branch relation reasoning to capture critical object relationships between the main branch and auxiliary branches. Moreover, we apply a diversity loss to make each branch only pay attention to its corresponding object and boost multi-branch learning. The extensive experiments show the effectiveness of our proposed method.

📄 PDF Abstract BibTeX arXiv:2008.06941

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityObjectRelationRelation NetworkSentenceSpatio-Temporal Video GroundingVideo Grounding

Similar Papers 제목 키워드 기반

GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object Tracking

2025-01-01 · CVPR 2025 1 · Hyunseop Kim, Hyo-Jun Lee, Yonguk Lee, Jinu Lee 외

Recently, 3D multi-object tracking (MOT) has widely adopted the standard tracking-by-detection paradigm, which solves the association problem between detections and tracks. Many tracking-by-detection approaches estab…

3D Multi-Object TrackingMulti-Object TrackingObject TrackingRelation

Relational Long Short-Term Memory for Video Action Recognition

2018-11-16 · Zexi Chen, Bharathkumar Ramachandra, Tianfu Wu, Ranga Raju Vatsavai

Spatial and temporal relationships, both short-range and long-range, between objects in videos, are key cues for recognizing actions. It is a challenging problem to model them jointly. In this paper, we first present a n…

Action RecognitionTemporal Action Localization

Spatio-Temporal Relation Learning for Video Anomaly Detection

2022-09-27 · Hui Lv, Zhen Cui, Biao Wang, Jian Yang

Anomaly identification is highly dependent on the relationship between the object and the scene, as different/same object actions in same/different scenes may lead to various degrees of normality and anomaly. Therefore, …

Anomaly DetectionGraph EmbeddingKnowledge Graph EmbeddingObject+4

Symbiotic Attention with Privileged Information for Egocentric Action Recognition

2020-02-08 · Xiaohan Wang, Yu Wu, Linchao Zhu, Yi Yang

Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recog…

Action RecognitionEgocentric Activity RecognitionGeneral Classificationobject-detection+3

Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering

2023-07-25 · Yi Cheng, Hehe Fan, Dongyun Lin, Ying Sun 외

The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually …

graph constructionQuestion AnsweringRelationVideo Question Answering