paper-with-me

홈 › Papers

A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving

2026-01-07 · Liangdong Zhang, Yiming Nie, Haoyang Li, Fanjie Kong, Baobao Zhang, Shunxin Huang, Kai Fu, Chen Min, Liang Xiao arxiv

Efficient trajectory planning in off-road terrains presents a formidable challenge for autonomous vehicles, often necessitating complex multi-step pipelines. However, traditional approaches exhibit limited adaptability in dynamic environments. To address these limitations, this paper proposes OFF-EMMA, a novel end-to-end multimodal framework designed to overcome the deficiencies of insufficient spatial perception and unstable reasoning in visual-language-action (VLA) models for off-road autonomous driving scenarios. The framework explicitly annotates input images through the design of a visual prompt block and introduces a chain-of-thought with self-consistency (COT-SC) reasoning strategy to enhance the accuracy and robustness of trajectory planning. The visual prompt block utilizes semantic segmentation masks as visual prompts, enhancing the spatial understanding ability of pre-trained visual-language models for complex terrains. The COT- SC strategy effectively mitigates the error impact of outliers on planning performance through a multi-path reasoning mechanism. Experimental results on the RELLIS-3D off-road dataset demonstrate that OFF-EMMA significantly outperforms existing methods, reducing the average L2 error of the Qwen backbone model by 13.3% and decreasing the failure rate from 16.52% to 6.56%.

📄 PDF Abstract BibTeX arXiv:2601.03519

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationAutonomous VehiclesTrajectory PlanningAutonomous Driving

Similar Papers 제목 키워드 기반

ROSGPT_Vision: Commanding Robots Using Only Language Models' Prompts

2023-08-22 · Bilel Benjdira, Anis Koubaa, Anas M. Ali

In this paper, we argue that the next generation of robots can be commanded using only Language Models' prompts. Every prompt interrogates separately a specific Robotic Modality via its Modality Language Model (MLM). A c…

Language ModelingLanguage ModellingLarge Language Model

Delving into Multimodal Prompting for Fine-grained Visual Classification

2023-09-16 · Xin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du 외

Fine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, preva…

ClassificationFine-Grained Image Classification

Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts

2024-06-04 · Haodong Hong, Sen Wang, Zi Huang, Qi Wu 외

Current Vision-and-Language Navigation (VLN) tasks mainly employ textual instructions to guide agents. However, being inherently abstract, the same textual instruction can be associated with different visual signals, cau…

NavigateVision and Language Navigation

RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System

2025-11-23 · Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningyuan Xiao 외 arxiv

Current roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introd…

Visual Question AnsweringComputational EfficiencyMulti-Task Learning

Out-of-Distribution Detection with Positive and Negative Prompt Supervision Using Large Language Models

2025-11-14 · Zhixia He, Chen Zhao, Minglai Shao, Xintao Wu 외 arxiv

Out-of-distribution (OOD) detection is committed to delineating the classification boundaries between in-distribution (ID) and OOD images. Recent advances in vision-language models (VLMs) have demonstrated remarkable OOD…

Out-of-Distribution Detection