paper-with-me

홈 › Papers

Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach

2025-07-27 · Yanming Xiu, Maria Gorlatova arxiv

The virtual content in augmented reality (AR) can introduce misleading or harmful information, leading to semantic misunderstandings or user errors. In this work, we focus on visual information manipulation (VIM) attacks in AR, where virtual content changes the meaning of real-world scenes in subtle but impactful ways. We introduce a taxonomy that categorizes these attacks into three formats: character, phrase, and pattern manipulation, and three purposes: information replacement, information obfuscation, and extra wrong information. Based on the taxonomy, we construct a dataset, AR-VIM, which consists of 452 raw-AR video pairs spanning 202 different scenes, each simulating a real-world AR scenario. To detect the attacks in the dataset, we propose a multimodal semantic reasoning framework, VIM-Sense. It combines the language and visual understanding capabilities of vision-language models (VLMs) with optical character recognition (OCR)-based textual analysis. VIM-Sense achieves an attack detection accuracy of 88.94% on AR-VIM, consistently outperforming vision-only and text-only baselines. The system achieves an average attack detection latency of 7.07 seconds in a simulated video processing framework and 7.17 seconds in a real-world evaluation conducted on a mobile Android AR application.

📄 PDF Abstract BibTeX arXiv:2507.20356

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality

2025-01-22 · Yanming Xiu, Tim Scargill, Maria Gorlatova

In Augmented Reality (AR), virtual content enhances user experience by providing additional information. However, improperly positioned or designed virtual content can be detrimental to task performance, as it can impair…

Language ModelingLanguage Modelling

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

2026-04-07 · Yanming Xiu, Zhengyuan Jiang, Neil Zhenqiang Gong, Maria Gorlatova arxiv

Augmented reality (AR) has rapidly expanded over the past decade. As AR becomes increasingly integrated into daily life, its security and reliability emerge as critical challenges. Among various threats, contradictory vi…

Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations

2025-09-16 · Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu 외 arxiv

The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifa…

Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models

2025-02-03 · Yuyang Gong, Zhuo Chen, Miaokun Chen, Fengchang Yu 외

Retrieval-Augmented Generation (RAG) systems based on Large Language Models (LLMs) have become essential for tasks such as question answering and content generation. However, their increasing impact on public opinion and…

Question AnsweringRAGRetrievalRetrieval-augmented Generation

SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors

2025-05-30 · Tianlong Yu, Chenghang Ye, Zheyu Yang, Ziyi Zhou 외

The SEAR Dataset is a novel multimodal resource designed to study the emerging threat of social engineering (SE) attacks orchestrated through augmented reality (AR) and multimodal large language models (LLMs). This datas…