paper-with-me

Papers

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

2026-04-07 · Yanming Xiu, Zhengyuan Jiang, Neil Zhenqiang Gong, Maria Gorlatova arxiv

Augmented reality (AR) has rapidly expanded over the past decade. As AR becomes increasingly integrated into daily life, its security and reliability emerge as critical challenges. Among various threats, contradictory virtual content attacks, where malicious or inconsistent virtual elements are introduced into the user's view, pose a unique risk by misleading users, creating semantic confusion, or delivering harmful information. In this work, we systematically model such attacks and present ContrAR, a novel benchmark for evaluating the robustness of vision-language models (VLMs) against virtual content manipulation and contradiction in AR. ContrAR contains 312 real-world AR videos validated by 10 human participants. We further benchmark 11 VLMs, including both commercial and open-source models. Experimental results reveal that while current VLMs exhibit reasonable understanding of contradictory virtual content, room still remains for improvement in detecting and reasoning about adversarial content manipulations in AR environments. Moreover, balancing detection accuracy and latency remains challenging.

📄 PDF Abstract BibTeX arXiv:2604.05510

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions

2024-08-02 · Jin Gao, Lei Gan, Yuankai Li, Yixin Ye 외

Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and context length, which is challenging fo…

Benchmarkingmultimodal interaction

SemEval-2021 Task 12: Learning with Disagreements

2021-08-01 · SEMEVAL 2021 · Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller 외

Disagreement between coders is ubiquitous in virtually all datasets annotated with human judgements in both natural language processing and computer vision. However, most supervised machine learning methods assume that a…

Benchmarking Visual-Inertial Deep Multimodal Fusion for Relative Pose Regression and Odometry-aided Absolute Pose Regression

2022-08-01 · Felix Ott, Nisha Lakshmana Raichur, David Rügamer, Tobias Feigl 외

Visual-inertial localization is a key problem in computer vision and robotics applications such as virtual reality, self-driving cars, and aerial vehicles. The goal is to estimate an accurate pose of an object when eithe…

BenchmarkingregressionSelf-Driving Cars

Benchmarking Language-agnostic Intent Classification for Virtual Assistant Platforms

2022-07-01 · NAACL (MIA) 2022 7 · Gengyu Wang, Cheng Qian, Lin Pan, Haode Qi 외

Current virtual assistant (VA) platforms are beholden to the limited number of languages they support. Every component, such as the tokenizer and intent classifier, is engineered for specific languages in these intricate…

BenchmarkingClassificationintent-classificationIntent Classification

Red Teaming Language Models for Processing Contradictory Dialogues

2024-05-16 · Xiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao Chen

Most language models currently available are prone to self-contradiction during dialogues. To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contrad…

Red Teamingvalid