paper-with-me

Papers

Position: Editing Large Language Models Poses Serious Safety Risks

2025-02-05 · Paul Youssef, Zhixue Zhao, Daniel Braun, Jörg Schlötterer, Christin Seifert

Large Language Models (LLMs) contain large amounts of facts about the world. These facts can become outdated over time, which has led to the development of knowledge editing methods (KEs) that can change specific facts in LLMs with limited side effects. This position paper argues that editing LLMs poses serious safety risks that have been largely overlooked. First, we note the fact that KEs are widely available, computationally inexpensive, highly performant, and stealthy makes them an attractive tool for malicious actors. Second, we discuss malicious use cases of KEs, showing how KEs can be easily adapted for a variety of malicious purposes. Third, we highlight vulnerabilities in the AI ecosystem that allow unrestricted uploading and downloading of updated models without verification. Fourth, we argue that a lack of social and institutional awareness exacerbates this risk, and discuss the implications for different stakeholders. We call on the community to (i) research tamper-resistant models and countermeasures against malicious model editing, and (ii) actively engage in securing the AI ecosystem.

📄 PDF Abstract BibTeX arXiv:2502.02958

Code (0)

등록된 구현이 없습니다.

Tasks

knowledge editingModel EditingPosition

Similar Papers 제목 키워드 기반

Knowledge in Superposition: Unveiling the Failures of Lifelong Knowledge Editing for Large Language Models

2024-08-14 · Chenhui Hu, Pengfei Cao, Yubo Chen, Kang Liu 외

Knowledge editing aims to update outdated or incorrect knowledge in large language models (LLMs). However, current knowledge editing methods have limited scalability for lifelong editing. This study explores the fundamen…

knowledge editing

ReLayout: Versatile and Structure-Preserving Design Layout Editing via Relation-Aware Design Reconstruction

2026-02-01 · Jiawei Lin, Shizhao Sun, Danqing Huang, Ting Liu 외 arxiv

Automated redesign without manual adjustments marks a key step forward in the design workflow. In this work, we focus on a foundational redesign task termed design layout editing, which seeks to autonomously modify the g…

Relation Also Knows: Rethinking the Recall and Editing of Factual Associations in Auto-Regressive Transformer Language Models

2024-08-27 · Xiyu Liu, Zhengxiao Liu, Naibin Gu, Zheng Lin 외

The storage and recall of factual associations in auto-regressive transformer language models (LMs) have drawn a great deal of attention, inspiring knowledge editing by directly modifying the located model weights. Most …

knowledge editingRelationSpecificity

GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing

2024-07-08 · Zhenyu Wang, Aoxue Li, Zhenguo Li, Xihui Liu

Despite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mec…

Image GenerationLanguage ModelingLanguage ModellingLarge Language Model+2

Untying the Reversal Curse via Bidirectional Language Model Editing

2023-10-16 · Jun-Yu Ma, Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 외

Recent studies have demonstrated that large language models (LLMs) store massive factual knowledge within their parameters. But existing LLMs are prone to hallucinate unintended text due to false or outdated knowledge. S…

knowledge editingLanguage ModelingLanguage ModellingModel Editing+1