paper-with-me

Papers

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

2025-05-23 · Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, Xing Wei

Visual language models (VLMs) have attracted increasing interest in autonomous driving due to their powerful reasoning capabilities. However, existing VLMs typically utilize discrete text Chain-of-Thought (CoT) tailored to the current scenario, which essentially represents highly abstract and symbolic compression of visual information, potentially leading to spatio-temporal relationship ambiguity and fine-grained information loss. Is autonomous driving better modeled on real-world simulation and imagination than on pure symbolic logic? In this paper, we propose a spatio-temporal CoT reasoning method that enables models to think visually. First, VLM serves as a world model to generate unified image frame for predicting future world states: where perception results (e.g., lane divider and 3D detection) represent the future spatial relationships, and ordinary future frame represent the temporal evolution relationships. This spatio-temporal CoT then serves as intermediate reasoning steps, enabling the VLM to function as an inverse dynamics model for trajectory planning based on current observations and future predictions. To implement visual generation in VLMs, we propose a unified pretraining paradigm integrating visual generation and understanding, along with a progressive visual CoT enhancing autoregressive image generation. Extensive experimental results demonstrate the effectiveness of the proposed method, advancing autonomous driving towards visual reasoning.

📄 PDF Abstract BibTeX arXiv:2505.17685

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingImage GenerationTrajectory PlanningVisual Reasoning

Similar Papers 제목 키워드 기반

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

2026-03-13 · Yuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu 외 arxiv

Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLL…

Visual Question Answering

LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving

2026-03-02 · Yuechen Luo, Fang Li, Shaoqing Xu, Yang Ji 외 arxiv

While Vision-Language-Action (VLA) models have revolutionized autonomous driving by unifying perception and planning, their reliance on explicit textual Chain-of-Thought (CoT) leads to semantic-perceptual decoupling and …

Reinforcement LearningAutonomous Driving

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

2026-08-21 · Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang 외 arxiv

Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion a…

Autonomous Driving

Deep Learning Based Motion Planning For Autonomous Vehicle Using Spatiotemporal LSTM Network

2019-03-05 · Zhengwei Bai, Baigen Cai, Wei Shangguan, Linguo Chai

Motion Planning, as a fundamental technology of automatic navigation for the autonomous vehicle, is still an open challenging issue in the real-life traffic situation and is mostly applied by the model-based approaches. …

Motion Planning

A Knowledge-Graph Translation Layer for Mission-Aware Multi-Agent Path Planning in Spatiotemporal Dynamics

2025-10-24 · Edward Holmberg, Elias Ioup, Mahdi Abdelguerfi arxiv

The coordination of autonomous agents in dynamic environments is hampered by the semantic gap between high-level mission objectives and low-level planner inputs. To address this, we introduce a framework centered on a Kn…