paper-with-me

홈 › Papers

Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming

2026-04-07 · Baoshun Tong, Haoran He, Ling Pan, Yang Liu, Liang Lin arxiv

Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to linguistic nuances remains a critical, under-explored safety concern, posing a significant safety risk to real-world deployment. Red teaming, or identifying environmental scenarios that elicit catastrophic behaviors, is an important step in ensuring the safe deployment of embodied AI agents. Reinforcement learning (RL) has emerged as a promising approach in automated red teaming that aims to uncover these vulnerabilities. However, standard RL-based adversaries often suffer from severe mode collapse due to their reward-maximizing nature, which tends to converge to a narrow set of trivial or repetitive failure patterns, failing to reveal the comprehensive landscape of meaningful risks. To bridge this gap, we propose a novel \textbf{D}iversity-\textbf{A}ware \textbf{E}mbodied \textbf{R}ed \textbf{T}eaming (\textbf{DAERT}) framework, to expose the vulnerabilities of VLAs against linguistic variations. Our design is based on evaluating a uniform policy, which is able to generate a diverse set of challenging instructions while ensuring its attack effectiveness, measured by execution failures in a physical simulator. We conduct extensive experiments across different robotic benchmarks against two state-of-the-art VLAs, including $π_0$ and OpenVLA. Our method consistently discovers a wider range of more effective adversarial instructions that reduce the average task success rate from 93.33\% to 5.85\%, demonstrating a scalable approach to stress-testing VLA agents and exposing critical safety blind spots before real-world deployment.

📄 PDF Abstract BibTeX arXiv:2604.05595

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningRed Teaming

Similar Papers 제목 키워드 기반

Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity

2025-07-30 · Xinwei Wu, Haojie Li, Hongyu Liu, Xinyu Ji 외 arxiv

In this work, we study a critical research problem regarding the trustworthiness of large language models (LLMs): how LLMs behave when encountering ambiguous narrative text, with a particular focus on Chinese textual amb…

Uncovering Constraint-Based Behavior in Neural Models via Targeted Fine-Tuning

2021-06-02 · ACL 2021 5 · Forrest Davis, Marten Van Schijndel

A growing body of literature has focused on detailing the linguistic knowledge embedded in large, pretrained language models. Existing work has shown that non-linguistic biases in models can drive model behavior away fro…

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning

2026-04-07 · Tillmann Rheude, Stefan Hegselmann, Roland Eils, Benjamin Wild arxiv

Contrastive learning has become a standard approach for unsupervised learning from paired data, as demonstrated by CLIP for image-text matching. However, many domains involve more than two modalities and require objectiv…

Cross-Modal RetrievalContrastive LearningImage-text matching

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

2026-07-03 · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue 외 arxiv

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vi…

multimodal generation

Uncovering Probabilistic Implications in Typological Knowledge Bases

2019-06-18 · ACL 2019 7 · Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell, Isabelle Augenstein

The study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have post-positions. Uncovering such implications…

Knowledge Base Population