paper-with-me

홈 › Papers

Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark

2025-09-27 · Xinjie Shen, Mufei Li, Pan Li arxiv

The deployment of Large Language Models (LLMs) in embodied agents creates an urgent need to measure their privacy awareness in the physical world. Existing evaluation methods, however, are confined to natural language based scenarios. To bridge this gap, we introduce EAPrivacy, a comprehensive evaluation benchmark designed to quantify the physical-world privacy awareness of LLM-powered agents. EAPrivacy utilizes procedurally generated scenarios across four tiers to test an agent's ability to handle sensitive objects, adapt to changing environments, balance task execution with privacy constraints, and resolve conflicts with social norms. Our measurements reveal a critical deficit in current models. The top-performing model, Gemini 2.5 Pro, achieved only 59\% accuracy in scenarios involving changing physical environments. Furthermore, when a task was accompanied by a privacy request, models prioritized completion over the constraint in up to 86\% of cases. In high-stakes situations pitting privacy against critical social norms, leading models like GPT-4o and Claude-3.5-haiku disregarded the social norm over 15\% of the time. These findings, demonstrated by our benchmark, underscore a fundamental misalignment in LLMs regarding physically grounded privacy and establish the need for more robust, physically-aware alignment. Codes and datasets will be available at https://github.com/Graph-COM/EAPrivacy.

📄 PDF Abstract BibTeX arXiv:2510.02356

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

2026-05-06 · Junran Wang, Xinjie Shen, Zehao Jin, Pan Li arxiv

As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, the…

Teaching Physical Awareness to LLMs through Sounds

2025-06-10 · Weiguo Wang, Andy Nie, Wenrui Zhou, Yi Kai 외

Large Language Models (LLMs) have shown remarkable capabilities in text and multimodal processing, yet they fundamentally lack physical awareness--understanding of real-world physical phenomena. In this work, we present …

Direction of Arrival Estimation

Measuring Fairness Under Unawareness of Sensitive Attributes: A Quantification-Based Approach

2021-09-17 · Alessandro Fabris, Andrea Esuli, Alejandro Moreo, Fabrizio Sebastiani

Algorithms and models are increasingly deployed to inform decisions about people, inevitably affecting their lives. As a consequence, those in charge of developing these models must carefully evaluate their impact on dif…

Fairness

DEAN: Deactivating the Coupled Neurons to Mitigate Fairness-Privacy Conflicts in Large Language Models

2024-10-22 · Chen Qian, Dongrui Liu, Jie Zhang, Yong liu 외

Ensuring awareness of fairness and privacy in Large Language Models (LLMs) is critical. Interestingly, we discover a counter-intuitive trade-off phenomenon that enhancing an LLM's privacy awareness through Supervised Fin…

Fairness

Active World-Model with 4D-informed Retrieval for Exploration and Awareness

2026-04-17 · Elaheh Vaezpour, Amirhosein Javadi, Tara Javidi arxiv

Physical awareness, especially in a large and dynamic environment, is shaped by sensing decisions that determine observability across space, time, and scale, while observations impact the quality of sensing decisions. Th…

Reinforcement Learning