paper-with-me

홈 › Papers

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

2026-01-25 · Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin arxiv

Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized agents prioritizes utility and user experience, treating memory as a neutral component and largely overlooking its safety implications. In this paper, we reveal intent legitimation, a previously underexplored safety failure in personalized agents, where benign personal memories bias intent inference and cause models to legitimize inherently harmful queries. To study this phenomenon, we introduce PS-Bench, a benchmark designed to identify and quantify intent legitimation in personalized interactions. Across multiple memory-augmented agent frameworks and base LLMs, personalization increases attack success rates by 15.8\%--243.7\% relative to stateless baselines. We further provide mechanistic evidence for intent legitimation from internal representations space, and propose a lightweight detection-reflection method that effectively reduces safety degradation. Overall, our work provides the first systematic exploration and evaluation of intent legitimation as a safety failure mode that naturally arises from benign, real-world personalization, highlighting the importance of assessing safety under long-term personal context. Our code is available at: https://github.com/MuyuenLP/PS-Bench. WARNING: This paper may contain harmful content.

📄 PDF Abstract BibTeX arXiv:2601.17887

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

2026-06-08 · Yanyan Luo, Xue Han, Ruiqiao Bai, Xin Huang 외 arxiv

Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. However, the mechanisms that enable personalization also expand the s…

Reinforcement Learning

PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

2026-02-01 · Sidharth Pulipaka, Oliver Chen, Manas Sharma, Taaha S Bajwa 외 arxiv

Conversational assistants are increasingly integrating long-term memory with large language models (LLMs). This persistence of memories, e.g., the user is vegetarian, can enhance personalization in future conversations. …

Uncovering Safety Risks of Large Language Models through Concept Activation Vector

2024-04-18 · Zhihao Xu, Ruixuan Huang, Changyu Chen, Xiting Wang

Despite careful safety alignment, current large language models (LLMs) remain vulnerable to various attacks. To further unveil the safety risks of LLMs, we introduce a Safety Concept Activation Vector (SCAV) framework, w…

Safety Alignment

AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

2024-07-11 · Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan 외

Foundation models (FMs) provide societal benefits but also amplify risks. Governments, companies, and researchers have proposed regulatory frameworks, acceptable use policies, and safety benchmarks in response. However, …

Common Sense Reasoning

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

2026-07-01 · Tong Xu, Xinzhe Cao, Zhihui Zhu, Keyan Ding 외 arxiv

Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, ma…