paper-with-me

홈 › Papers

ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving

2026-04-07 · Kaiser Hamid, Can Cui, Nade Liang arxiv

Recent progress in vision-language-action (VLA) models has enabled language-conditioned driving agents to execute natural-language navigation commands in closed-loop simulation, yet standard evaluations largely assume instructions are precise and well-formed. In deployment, instructions vary in phrasing and specificity, may omit critical qualifiers, and can occasionally include misleading, authority-framed text, leaving instruction-level robustness under-measured. We introduce ICR-Drive, a diagnostic framework for instruction counterfactual robustness in end-to-end language-conditioned autonomous driving. ICR-Drive generates controlled instruction variants spanning four perturbation families: Paraphrase, Ambiguity, Noise, and Misleading, where Misleading variants conflict with the navigation goal and attempt to override intent. We replay identical CARLA routes under matched simulator configurations and seeds to isolate performance changes attributable to instruction language. Robustness is quantified using standard CARLA Leaderboard metrics and per-family performance degradation relative to the baseline instruction. Experiments on LMDrive and BEVDriver show that minor instruction changes can induce substantial performance drops and distinct failure modes, revealing a reliability gap for deploying embodied foundation models in safety-critical driving.

📄 PDF Abstract BibTeX arXiv:2604.05378

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data

2025-04-16 · Suyoung Bae, Hyojun Kim, YunSeok Choi, Jee-Hyong Lee

In various natural language processing (NLP) tasks, fine-tuning Pre-trained Language Models (PLMs) often leads to the issue of spurious correlations, which negatively impacts performance, particularly when dealing with o…

Contrastive LearningcounterfactualNatural Language InferenceSentence+2

Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Framework

2026-03-08 · Kaihua Tang, Jiaxin Qi, Jinli Ou, Yuhua Zheng 외 arxiv

The emergence of Large Language Models (LLMs) has driven rapid progress in multi-modal learning, particularly in the development of Large Vision-Language Models (LVLMs). However, existing LVLM training paradigms place ex…

RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos

2025-10-10 · Zixi Yang, Jiapeng Li, Muxi Diao, Yinuo Jing 외 arxiv

Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when faced with manipulated video content, r…

Counterfactual Vision-and-Language Navigation via Adversarial Path Sampler

2020-08-01 · ECCV 2020 8 · Tsu-Jui Fu, Xin Eric Wang, Matthew F. Peterson,Scott T. Grafton, Miguel P. Eckstein 외

Vision-and-Language Navigation (VLN) is a task where agents must decide how to move through a 3D environment to reach a goal by grounding natural language instructions to the visual surroundings. One of the problems of t…

counterfactualCounterfactual ReasoningData AugmentationVision and Language Navigation

Counterfactual Vision-and-Language Navigation via Adversarial Path Sampling

2019-11-17 · Tsu-Jui Fu, Xin Eric Wang, Matthew Peterson, Scott Grafton 외

Vision-and-Language Navigation (VLN) is a task where agents must decide how to move through a 3D environment to reach a goal by grounding natural language instructions to the visual surroundings. One of the problems of t…

counterfactualCounterfactual ReasoningData AugmentationVision and Language Navigation