paper-with-me

홈 › Papers

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

2026-07-02 · Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma arxiv

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world applications in specialized fields, where models inevitably encounter rare visual concepts and complex spatio-temporal dynamics. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured videos such as expert-annotated mouse behaviors with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated training subsets to systematically measure domain adaptability. We extensively evaluate 15 state-of-the-art VLMs, assessing their zero-shot generalization and In-Context Learning (ICL) capabilities under practical computational constraints. Ultimately, our findings reveal that current models fail in both zero-shot and ICL-based adaptation when confronted with specialized domains, exposing critical flaws in spatio-temporal reasoning that future research must address.

📄 PDF Abstract BibTeX arXiv:2607.02269

Code (0)

등록된 구현이 없습니다.

Tasks

Spatio-Temporal Video GroundingZero-shot GeneralizationDomain Adaptation

Similar Papers 제목 키워드 기반

Conditional Multi-Event Temporal Grounding in Long-Form Video

2026-06-13 · Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez, Lincoln Spencer 외 arxiv

Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositional temporal and spatial conditions. Ex…

VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks

2025-12-18 · Beitong Zhou, Zhexiao Huang, Yuan Guo, Zhangxuan Gu 외 arxiv

GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, o…

Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization

2026-04-14 · Zanyi Wang, Fan Li, Dengyang Jiang, Liuzhuozheng Li 외 arxiv

Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notoriously data-hungry. However, gathering large-scale STVG data is exceptio…

Spatio-Temporal Video Grounding

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

2025-10-19 · Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa 외 arxiv

We introduce ED-VTG, a method for fine-grained video temporal grounding utilizing multi-modal large language models. Our approach harnesses the capabilities of multimodal LLMs to jointly process text and video, in order …

Natural Language QueriesVideo Grounding

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

2025-07-06 · Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian 외 arxiv

With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional domains such as medical education. However,…

Information Retrieval