paper-with-me

Video Visual Relation Detection

2개 벤치마크 · 논문 16편 · 이 태스크의 논문 보기 →

Benchmarks

ImageNet-VidVRD

결과 8개

VidOR

결과 8개

Most implemented

Papers

Spatial-Temporal Human-Object Interaction Detection

2025-08-24 · Xu Sun, Yunqing He, Tongwei Ren, Gangshan Wu arxiv

In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects an…

Human-Object Interaction DetectionVideo Visual Relation Detection

OpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment

2025-03-12 · Qi Liu, Weiying Xue, Yuxiao Wang, Zhenao Wei

The video visual relation detection (VidVRD) task is to identify objects and their relationships in videos, which is challenging due to the dynamic content, high annotation costs, and long-tailed distribution of relation…

Prompt LearningRelationVideo Visual Relation Detection

VrdONE: One-stage Video Visual Relation Detection

2024-08-18 · Xinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu 외

Video Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks. Traditional metho…

Predicate DetectionRelationVideo Visual Relation Detection

SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos

2024-04-06 · CVPR 2024 1 · Tao Wu, Runyu He, Gangshan Wu, LiMin Wang

Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitat…

Graph GenerationRelationScene Graph GenerationVideo scene graph generation+2

In Defense of Clip-based Video Relation Detection

2023-07-18 · Meng Wei, Long Chen, Wei Ji, Xiaoyu Yue 외

Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be broadly categorized into bottom-up and t…

Feature CompressionObject TrackingRelationVideo Visual Relation Detection

Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

2023-02-01 · Kaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao 외

Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary predictions trained on limited base categories, e.g., object classification and detection. In this paper, we propose compositional…

ObjectRelationVideo Visual Relation Detection

전체 16편 보기 →