Video Visual Relation Detection
2개 벤치마크 · 논문 16편 · 이 태스크의 논문 보기 →
Benchmarks
ImageNet-VidVRD
VidOR
Most implemented
Spatial-Temporal Transformer for Dynamic Scene Graph Generation
VrdONE: One-stage Video Visual Relation Detection
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Video Relation Detection via Tracklet based Visual Transformer
Social Fabric: Tubelet Compositions for Video Relation Detection
Papers
Spatial-Temporal Human-Object Interaction Detection
In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects an…
Human-Object Interaction DetectionVideo Visual Relation DetectionOpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment
The video visual relation detection (VidVRD) task is to identify objects and their relationships in videos, which is challenging due to the dynamic content, high annotation costs, and long-tailed distribution of relation…
Prompt LearningRelationVideo Visual Relation DetectionVrdONE: One-stage Video Visual Relation Detection
Video Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks. Traditional metho…
Predicate DetectionRelationVideo Visual Relation DetectionSportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitat…
Graph GenerationRelationScene Graph GenerationVideo scene graph generation+2In Defense of Clip-based Video Relation Detection
Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be broadly categorized into bottom-up and t…
Feature CompressionObject TrackingRelationVideo Visual Relation DetectionCompositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary predictions trained on limited base categories, e.g., object classification and detection. In this paper, we propose compositional…
ObjectRelationVideo Visual Relation Detection