paper-with-me

홈 › Papers

ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models

2025-03-25 · Dohwan Ko, Sihyeon Kim, Yumin Suh, Vijay Kumar B. G, Minseo Yoon, Manmohan Chandraker, Hyunwoo J. Kim

Spatio-temporal reasoning is essential in understanding real-world environments in various fields, eg, autonomous driving and sports analytics. Recent advances have improved the spatial reasoning ability of Vision-Language Models (VLMs) by introducing large-scale data, but these models still struggle to analyze kinematic elements like traveled distance and speed of moving objects. To bridge this gap, we construct a spatio-temporal reasoning dataset and benchmark involving kinematic instruction tuning, referred to as STKit and STKit-Bench. They consist of real-world videos with 3D annotations, detailing object motion dynamics: traveled distance, speed, movement direction, inter-object distance comparisons, and relative movement direction. To further scale such data construction to videos without 3D labels, we propose an automatic pipeline to generate pseudo-labels using 4D reconstruction in real-world scale. With our kinematic instruction tuning data for spatio-temporal reasoning, we present ST-VLM, a VLM enhanced for spatio-temporal reasoning, which exhibits outstanding performance on STKit-Bench. Furthermore, we show that ST-VLM generalizes robustly across diverse domains and tasks, outperforming baselines on other spatio-temporal benchmarks (eg, ActivityNet, TVQA+). Finally, by integrating learned spatio-temporal reasoning with existing abilities, ST-VLM enables complex multi-step reasoning. Project page: https://ikodoh.github.io/ST-VLM.

📄 PDF Abstract BibTeX arXiv:2503.19355

Code (0)

등록된 구현이 없습니다.

Tasks

4D reconstructionAutonomous DrivingSpatial ReasoningSports Analytics

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data

2025-09-03 · Honglu Zhou, Xiangyu Peng, Shrikant Kendre, Michael S. Ryoo 외 arxiv

Next-generation AI companions must go beyond general video understanding to resolve spatial and temporal references in dynamic, real-world environments. Existing Video Large Language Models (Video LLMs), while capable of…

MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence

2026-02-28 · Xingyilang Yin, Chengzhengxu Li, Jiahao Chang, Chi-Man Pun 외 arxiv

Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability rem…

CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks

2025-07-18 · Yanan Wang, Julio Vizcarra, Zhi Li, Hao Niu 외 arxiv

Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fine-grained object-level video understand…

Temporal Relation Extraction

TrajGenAgent: A Hierarchical LLM Agent for Human Mobility Trajectory Generation

2026-06-10 · Siyu Li, Toan Tran, Lingyi Zhao, Khurram Shafique 외 arxiv

Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating realistic synthetic trajectory generati…

Prompt Engineering

ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation

2026-04-20 · Chuanhao Ma, Hanyu Zhou, Shihan Peng, Yan Li 외 arxiv

Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal manipulation. Typically, existing methods mainly embed spatiotemporal kno…