paper-with-me

Papers

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

2025-03-31 · Yun Li, Yiming Zhang, Tao Lin, Xiangrui Liu, Wenxiao Cai, Zheng Liu, Bo Zhao

The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their ability to perform precise and quantitative spatial-temporal understanding in real-world applications remains largely unexamined, leading to uncertain prospects. To evaluate models' Spatial-Temporal Intelligence, we introduce STI-Bench, a benchmark designed to evaluate MLLMs' spatial-temporal understanding through challenging tasks such as estimating and predicting the appearance, pose, displacement, and motion of objects. Our benchmark encompasses a wide range of robot and vehicle operations across desktop, indoor, and outdoor scenarios. The extensive experiments reveals that the state-of-the-art MLLMs still struggle in real-world spatial-temporal understanding, especially in tasks requiring precise distance estimation and motion analysis.

📄 PDF Abstract BibTeX arXiv:2503.23765

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

2026-05-08 · Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami arxiv

Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorly understood. Existing multi-video bench…

Spatial Reasoning

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

2026-05-21 · Jinho Park, Youbin Kim, Hogun Park, Eunbyung Park arxiv

Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-tempor…

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

2026-07-23 · Han Li, Si Liu, Zehao Huang, Dongxin Lyu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation…

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

2025-05-17 · Zihao Dongfang, Xu Zheng, Ziqiao Weng, Yuanhuiyi Lyu 외

The 180x360 omnidirectional field of view captured by 360-degree cameras enables their use in a wide range of applications such as embodied AI and virtual reality. Although recent advances in multimodal large language mo…

HallucinationObject CountingSpatial Reasoning

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

2025-09-23 · Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato 외 arxiv

Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely o…