paper-with-me

홈 › Papers

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

2026-06-27 · Tianshu Zhang, Yan Wang, Ji Qi, Lijie Wen arxiv

Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language models (VLMs) show strong reasoning ability, directly applying frame-by-frame inference to long sequences is computationally expensive and unstable. We propose a practical pipeline that shifts from frame-level to second-level tracking and performs cross-second smoothing to preserve continuity while reducing sequence length. To improve reasoning supervision, we synthesize chain-of-thought style trajectories using advanced multimodal models for temporal localization and target selection, and replace generated spatio-temporal coordinates with ground-truth annotations to avoid noisy supervision. We further optimize the policy with reinforcement learning using a verifier based on $t\_\mathrm{IoU}+mv\_\mathrm{IoU}$. Experiments across multiple FPS settings show that our method achieves a strong trade-off between efficiency and localization quality.

📄 PDF Abstract BibTeX arXiv:2606.29023

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningObject Tracking

Similar Papers 제목 키워드 기반

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

2025-03-18 · Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li 외

Multimodal large language models (MLLMs) have made remarkable progress in either temporal or spatial localization. However, they struggle to perform spatio-temporal video grounding. This limitation stems from two major c…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

2025-09-18 · Zaiquan Yang, Yuhao Liu, Gerhard Hancke, Rynson W. H. Lau arxiv

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large language models (MLLMs) to explore a zero-sh…

Spatio-Temporal Video Grounding

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

2025-11-04 · Jinhwan Seo, Yoonki Cho, Junhyug Noh, Sung-eui Yoon arxiv

In this technical report, we introduce a framework to address Grounded Video Question Answering (GVQA) task for the ICCV 2025 Perception Test Challenge. The GVQA task demands robust multimodal models capable of complex r…

Video Question Answering

What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

2024-01-01 · CVPR 2024 1 · Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann 외

Spatio-temporal grounding describes the task of localizing events in space and time e.g. in video data based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bo…

Representation Learning

What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

2023-03-29 · Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann 외

Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bou…

Representation LearningSpatio-Temporal Video Grounding