paper-with-me

홈 › Papers

doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation

2024-12-08 · Parthib Roy, Srinivasa Perisetla, Shashank Shriram, Harsha Krishnaswamy, Aryan Keskar, Ross Greer

Human-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning. This paper introduces doScenes, a novel dataset designed to facilitate research on human-vehicle instruction interactions, focusing on short-term directives that directly influence vehicle motion. By annotating multimodal sensor data with natural language instructions and referentiality tags, doScenes bridges the gap between instruction and driving response, enabling context-aware and adaptive planning. Unlike existing datasets that focus on ranking or scene-level reasoning, doScenes emphasizes actionable directives tied to static and dynamic scene objects. This framework addresses limitations in prior research, such as reliance on simulated data or predefined action sets, by supporting nuanced and flexible responses in real-world scenarios. This work lays the foundation for developing learning strategies that seamlessly integrate human instructions into autonomous systems, advancing safe and effective human-vehicle collaboration for vision-language navigation. We make our data publicly available at https://www.github.com/rossgreer/doScenes

📄 PDF Abstract BibTeX arXiv:2412.05893

Code (1)

rossgreer/doscenes 공식 구현

Tasks

Autonomous DrivingAutonomous VehiclesMotion PlanningVision-Language Navigation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A DVDrive Approach for doScenes Instructed Driving Challenge

2026-06-19 · Zijian Fu, Xiangyang Chu, Mengshi Qi, Huadong Ma 외 arxiv

Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natur…

Trajectory PredictionAutonomous DrivingVisual Grounding

Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models

2026-02-04 · Angel Martinez-Sanchez, Parthib Roy, Ross Greer arxiv

Instruction-grounded driving, where passenger language guides trajectory planning, requires vehicles to understand intent before motion. However, most prior instruction-following planners rely on simulation or fixed comm…

Trajectory PlanningAutonomous DrivingMotion Planning

Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning

2026-02-07 · Ross Greer, Maitrayee Keskar, Angel Martinez-Sanchez, Parthib Roy 외 arxiv

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-c…

Visual Question AnsweringRepresentation LearningTrajectory PlanningAutonomous Driving

A Comprehensive LLM-powered Framework for Driving Intelligence Evaluation

2025-03-07 · Shanhe You, Xuewen Luo, Xinhe Liang, Jiashu Yu 외

Evaluation methods for autonomous driving are crucial for algorithm optimization. However, due to the complexity of driving intelligence, there is currently no comprehensive evaluation method for the level of autonomous …

Autonomous Driving

An interactive enhanced driving dataset for autonomous driving

2026-02-24 · Haojie Feng, Peizhi Zhang, Mengjie Tian, Xinrui Zhang 외 arxiv

The evolution of autonomous driving towards full automation demands robust interactive capabilities; however, the development of Vision-Language-Action (VLA) models is constrained by the sparsity of interactive scenarios…

Autonomous Driving