paper-with-me

Video Grounding

2개 벤치마크 · 논문 140편 · 이 태스크의 논문 보기 →

Benchmarks

QVHighlights

결과 21개

MAD

결과 6개

Most implemented

Papers

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

2026-08-24 · Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu 외 arxiv

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active tem…

Reinforcement LearningVideo Grounding

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

2026-06-15 · Ke Li, Di Wang, Yongshan Zhu, Ting Wang 외 arxiv

Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotati…

Referring ExpressionVisual GroundingVideo Grounding

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

2026-05-31 · Xin Dong, Wenjia Geng, Wenfeng Deng, Yansong Tang arxiv

Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat the…

Representation LearningHighlight DetectionMoment RetrievalVideo Grounding

Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video

2026-05-02 · Jiantang Huang arxiv

Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate joint localization in time, space, and collision type is required. We pr…

Video Grounding

Static and Dynamic Graph Alignment Network for Temporal Video Grounding

2026-05-01 · Zhanjie Hu, Bolin Zhang, Jianhua Wang, Jianbo Zheng 외 arxiv

Temporal Video Grounding (TVG) aims to localize temporal moments in an untrimmed video that semantically correspond to given natural language queries. Recently, Graph Convolutional Networks (GCN) have been widely adopted…

Natural Language QueriesContrastive LearningVideo Grounding

Subjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing

2026-04-27 · Cheng-Han Lee, Maniratnam Mandal, Neil Birkbeck, Yilin Wang 외 arxiv

With the rise of mobile video consumption on diverse handheld display resolutions and orientation modes, altering videos to aspect ratios poses challenges. Static cropping and border padding often compromises visual qual…

Video Grounding

전체 140편 보기 →