Video Grounding
2개 벤치마크 · 논문 140편 · 이 태스크의 논문 보기 →
Benchmarks
QVHighlights
MAD
Most implemented
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Context-Guided Spatio-Temporal Video Grounding
Papers
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active tem…
Reinforcement LearningVideo GroundingTraining-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos
Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotati…
Referring ExpressionVisual GroundingVideo GroundingCoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat the…
Representation LearningHighlight DetectionMoment RetrievalVideo GroundingTwo-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video
Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate joint localization in time, space, and collision type is required. We pr…
Video GroundingStatic and Dynamic Graph Alignment Network for Temporal Video Grounding
Temporal Video Grounding (TVG) aims to localize temporal moments in an untrimmed video that semantically correspond to given natural language queries. Recently, Graph Convolutional Networks (GCN) have been widely adopted…
Natural Language QueriesContrastive LearningVideo GroundingSubjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing
With the rise of mobile video consumption on diverse handheld display resolutions and orientation modes, altering videos to aspect ratios poses challenges. Static cropping and border padding often compromises visual qual…
Video Grounding