paper-with-me

Papers

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

2025-08-27 · Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu arxiv

Smartphones bring significant convenience to users but also enable devices to extensively record various types of personal information. Existing smartphone agents powered by Multimodal Large Language Models (MLLMs) have achieved remarkable performance in automating different tasks. However, as the cost, these agents are granted substantial access to sensitive users' personal information during this operation. To gain a thorough understanding of the privacy awareness of these agents, we present the first large-scale benchmark encompassing 7,138 scenarios to the best of our knowledge. In addition, for privacy context in scenarios, we annotate its type (e.g., Account Credentials), sensitivity level, and location. We then carefully benchmark seven available mainstream smartphone agents. Our results demonstrate that almost all benchmarked agents show unsatisfying privacy awareness (RA), with performance remaining below 60% even with explicit hints. Overall, closed-source agents show better privacy ability than open-source ones, and Gemini 2.0-flash achieves the best, achieving an RA of 67%. We also find that the agents' privacy detection capability is highly related to scenario sensitivity level, i.e., the scenario with a higher sensitivity level is typically more identifiable. We hope the findings enlighten the research community to rethink the unbalanced utility-privacy tradeoff about smartphone agents. Our code and benchmark are available at https://zhixin-l.github.io/SAPA-Bench.

📄 PDF Abstract BibTeX arXiv:2508.19493

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

2025-02-16 · Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang 외

Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a signific…

Benchmarking

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs

2025-09-06 · Zhaoyu Fan, Kaihang Pan, Mingze Zhou, Bosheng Qin 외 arxiv

Knowledge editing enables multimodal large language models (MLLMs) to efficiently update outdated or incorrect information. However, existing benchmarks primarily emphasize cognitive-level modifications while lacking a f…

knowledge editing

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

2025-03-28 · YuXuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads 외

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D datase…

BenchmarkingQuestion AnsweringVideo Question Answering

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

2026-03-23 · Qihui Zhu, Shouwei Ruan, Xiao Yang, Hao Jiang 외 arxiv

Despite the widespread adoption of MLLMs in embodied agents, their capabilities remain largely confined to reactive planning from immediate observations, consistently failing in spatial reasoning across extensive spatiot…

Spatial Reasoning

Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety

2025-08-29 · Younggun Kim, Sirnam Swetha, Fazil Kagdi, Mubarak Shah arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks. However, these models often infer and reveal sensitive biometric attributes such as race, gender, age, body wei…