Papers Video Visual Relation Detection
“Video Visual Relation Detection” 태그가 달린 논문 16편 · 필터 해제
Spatial-Temporal Human-Object Interaction Detection
In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects an…
Human-Object Interaction DetectionVideo Visual Relation DetectionOpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment
The video visual relation detection (VidVRD) task is to identify objects and their relationships in videos, which is challenging due to the dynamic content, high annotation costs, and long-tailed distribution of relation…
Prompt LearningRelationVideo Visual Relation DetectionVrdONE: One-stage Video Visual Relation Detection
Video Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks. Traditional metho…
Predicate DetectionRelationVideo Visual Relation DetectionSportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitat…
Graph GenerationRelationScene Graph GenerationVideo scene graph generation+2In Defense of Clip-based Video Relation Detection
Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be broadly categorized into bottom-up and t…
Feature CompressionObject TrackingRelationVideo Visual Relation DetectionCompositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary predictions trained on limited base categories, e.g., object classification and detection. In this paper, we propose compositional…
ObjectRelationVideo Visual Relation DetectionMeta Spatio-Temporal Debiasing for Video Scene Graph Generation
Video scene graph generation (VidSGG) aims to parse the video content into scene graphs, which involves modeling the spatio-temporal contextual information in the video. However, due to the long-tailed training data in d…
Graph GenerationMeta-LearningScene Graph GenerationVideo Visual Relation DetectionVRDFormer: End-to-End Video Visual Relation Detection With Transformers
Visual relation understanding plays an essential role for holistic video understanding. Most previous works adopt a multi-stage framework for video visual relation detection (VidVRD), which cannot capture long-term s…
ObjectRelationRelation ClassificationVideo Understanding+1Video Relation Detection via Tracklet based Visual Transformer
Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet detection pipeline MEGA and deepSORT to …
DecoderRelationVideo Visual Relation DetectionSocial Fabric: Tubelet Compositions for Video Relation Detection
This paper strives to classify and detect the relationship between object tubelets appearing within a video as a <subject-predicate-object> triplet. Where existing works treat object proposals or tubelets as single entit…
ObjectRelationTripletVideo Visual Relation Detection+1Spatial-Temporal Transformer for Dynamic Scene Graph Generation
Dynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects …
DecoderScene Graph GenerationVideo UnderstandingVideo Visual Relation Detection+1What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
Identifying relations between objects is central to understanding the scene. While several works have been proposed for relation modeling in the image domain, there have been many constraints in the video domain due to c…
RelationVideo Visual Relation DetectionVideo Visual Relation TaggingVideo Relation Detection with Trajectory-aware Multi-modal Features
Video relation detection problem refers to the detection of the relationship between different objects in videos, such as spatial relationship and action relationship. In this paper, we present video relation detection w…
Objectobject-detectionObject DetectionRelation+2LIGHTEN: Learning Interactions with Graph and Hierarchical TEmporal Networks for HOI in videos
Analyzing the interactions between humans and objects from a video includes identification of the relationships between humans and the objects present in the video. It can be thought of as a specialized version of Visual…
Human-Object Interaction DetectionRelationship DetectionVideo Visual Relation DetectionVisual Relationship DetectionBeyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global Context
Video visual relation detection (VidVRD) aims to describe all interacting objects in a video. Different from relationships in static images, videos contain an addition temporal channel. A majority of existing works divid…
RelationVideo Visual Relation DetectionVideo Relationship Reasoning using Gated Spatio-Temporal Energy Graph
Visual relationship reasoning is a crucial yet challenging task for understanding rich interactions across visual concepts. For example, a relationship 'man, open, door' involves a complex relation 'open' between concret…
Model OptimizationVideo RelationshipVideo Visual Relation Detection