paper-with-me

홈 › Papers

nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

2026-05-29 · Zhiyu Huang, Johnson Liu, Rui Song, Zewei Zhou, Ruining Yang, Yun Zhang, Tianhui Cai, Hanyin Zhang, Mingxuan Gao, Valeria Xu, Jiali Chen, Yishan Shen, Yiluan Guo, Tony, Qi, Jiaqi Ma arxiv

Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks mainly target perception, prediction, or planning, and provide limited supervision for reasoning over realistic long-tail driving scenes. We introduce nuReasoning, a large-scale real-world dataset and benchmark for reasoning-centric AD. Following the lineage of nuScenes and nuPlan, nuReasoning advances real-world AD datasets and benchmarks toward reasoning in long-tail driving scenarios. The dataset contains 20,000 clips, each 20 seconds long, collected across multiple cities, with synchronized multi-camera images, LiDAR data, HD maps, object annotations, and human-verified reasoning annotations spanning Spatial Reasoning, Decision Reasoning, and Counterfactual Reasoning. Unlike prior datasets that focus primarily on visual question answering, nuReasoning supports both reasoning evaluation and planning evaluation, enabling a direct study of how reasoning supervision affects driving performance. Experiments show that fine-tuning VLMs on nuReasoning substantially improves driving-specific question answering, while incorporating reasoning supervision into VLA training improves planning performance even when textual reasoning outputs are disabled at inference time. These results establish nuReasoning as a foundation for evaluating and improving robust, interpretable, reasoning-driven AD systems in realistic long-tail settings.

📄 PDF Abstract BibTeX arXiv:2605.31572

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringAutonomous DrivingSpatial Reasoning

Similar Papers 제목 키워드 기반

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

2026-08-13 · Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu 외 arxiv

Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events…

Video Retrieval

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

2026-06-13 · Yun Wang, Junbin Xiao, Han Lyu, Yifan Wang 외 arxiv

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence in egocentric video streams. UCS-Bench t…

Spatial Reasoning

Spatial-Conditioned Reasoning in Long-Egocentric Videos

2026-01-26 · James Tribble, Hao Wang, Si-En Hong, Chaoyi Zhou 외 arxiv

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and…

Spatial ReasoningVisual Navigation

K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

2026-07-02 · Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani 외 arxiv

MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies in understanding and modeling animal-cent…

Multimodal Reasoning

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

2026-03-10 · Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu 외 arxiv

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from …