Video scene graph generation
1개 벤치마크 · 논문 23편 · 이 태스크의 논문 보기 →
Benchmarks
ImageNet-VidVRD
Most implemented
Panoptic Video Scene Graph Generation
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
Papers
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models
Scene graph generation provides a compact structured representation for visual perception, but accurate and fast graph prediction from images and videos remains challenging. Recent VLM-based methods can generate scene gr…
Video scene graph generationReinforcement LearningFrequency-guided Multi-level Reasoning for Scene Graph Generation in Video
Video Scene Graph Generation aims to obtain structured semantic representations of objects and their relationships in videos for high-level understanding. However, existing methods still have limitations in handling long…
Video scene graph generationRelation ClassificationRevisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
Weakly-supervised video scene graph generation (WS-VSGG) aims to parse video content into structured relational triplets without bounding box annotations and with only sparse temporal labeling, significantly reducing ann…
Video scene graph generationSynthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
We introduce Synthetic Visual Genome 2 (SVG2), a large-scale panoptic video scene graph dataset. SVG2 contains over 636K videos with 6.6M objects, 52.0M attributes, and 6.7M relations, providing an order-of-magnitude inc…
Video scene graph generationVideo Question AnsweringPanoptic SegmentationSemantic ParsingClick2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
State-of-the-art Video Scene Graph Generation (VSGG) systems provide structured visual understanding but operate as closed, feed-forward pipelines with no ability to incorporate human guidance. In contrast, promptable se…
Video scene graph generationRelational ReasoningScene UnderstandingUNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typically target either coarse-grained box-…
Video scene graph generationRepresentation Learning