paper-with-me

Papers Model Editing

“Model Editing” 태그가 달린 논문 193편 · 필터 해제

Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm

2025-06-25 · Baixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu 외

Agents based on Large Language Models (LLMs) have demonstrated strong capabilities across a wide range of tasks. However, deploying LLM-based agents in high-stakes domains comes with significant safety and ethical risks.…

Model Editing

Mitigating Safety Fallback in Editing-based Backdoor Injection on LLMs

2025-06-16 · Houcheng Jiang, Zetong Zhao, Junfeng Fang, Haokai Ma 외

Large language models (LLMs) have shown strong performance across natural language tasks, but remain vulnerable to backdoor attacks. Recent model editing-based approaches enable efficient backdoor injection by directly m…

DiversityModel EditingSafety Alignment

DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models

2025-06-16 · Zhiyi Shi, Binjie Wang, Chongjie Si, Yichen Wu 외

Model editing aims to efficiently update a pre-trained model's knowledge without the need for time-consuming full retraining. While existing pioneering editing methods achieve promising results, they primarily focus on e…

Model Editing

SoK: Machine Unlearning for Large Language Models

2025-06-10 · Jie Ren, Yue Xing, Yingqian Cui, Charu C. Aggarwal 외

Large language model (LLM) unlearning has become a critical topic in machine learning, aiming to eliminate the influence of specific training data or knowledge without retraining the model from scratch. A variety of tech…

Large Language ModelMachine UnlearningModel Editing

MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs

2025-06-09 · Ke Wang, Yiming Qin, Nikolaos Dimitriadis, Alessandro Favero 외

Language models deployed in real-world systems often require post-hoc updates to incorporate new or corrected knowledge. However, editing such models efficiently and reliably - without retraining or forgetting previous i…

HallucinationModel EditingOut-of-Distribution GeneralizationQuestion Answering

The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing

2025-06-07 · Adrià Molina Rodríguez, Oriol Ramos Terrades, Josep Lladós

Achieving robustness in recognition systems across diverse domains is crucial for their practical utility. While ample data availability is usually assumed, low-resource languages, such as ancient manuscripts and non-wes…

Meta-LearningModel EditingOptical Character Recognition (OCR)Transfer Learning

Drop Dropout on Single-Epoch Language Model Pretraining

2025-05-30 · Houjun Liu, John Bauer, Christopher D. Manning

Originally, dropout was seen as a breakthrough regularization technique that reduced overfitting and improved performance in almost all applications of deep learning by reducing overfitting. Yet, single-epoch pretraining…

Language ModelingLanguage ModellingModel EditingNatural Language Inference+1

On Fairness of Task Arithmetic: The Role of Task Vectors

2025-05-30 · Hiroki Naganuma, Kotaro Yoshida, Laura Gomezjurado Gonzalez, Takafumi Horie 외

Model editing techniques, particularly task arithmetic using task vectors, have shown promise in efficiently modifying pre-trained models through arithmetic operations like task addition and negation. Despite computation…

FairnessHate Speech DetectionModel EditingNegation+2

Model Unlearning via Sparse Autoencoder Subspace Guided Projections

2025-05-30 · Xu Wang, Zihao Li, Benyou Wang, Yan Hu 외

Large language models (LLMs) store vast amounts of information, making them powerful yet raising privacy and safety concerns when selective knowledge removal is required. Existing unlearning strategies, ranging from grad…

Adversarial Robustnessfeature selectionGSM8KMMLU+2

DocMEdit: Towards Document-Level Model Editing

2025-05-26 · Li Zeng, Zeming Liu, Chong Feng, Heyan Huang 외

Model editing aims to correct errors and outdated knowledge in the Large language models (LLMs) with minimal cost. Prior research has proposed a variety of datasets to assess the effectiveness of these model editing meth…

modelModel Editing

REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing

2025-05-25 · Haitian Zhong, Yuhuan Liu, Ziyang Xu, Guofan Liu 외

Large language model editing methods frequently suffer from overfitting, wherein factual updates can propagate beyond their intended scope, overemphasizing the edited target even when it's contextually inappropriate. To …

knowledge editingLanguage ModelingLanguage ModellingLarge Language Model+1

Localizing Knowledge in Diffusion Transformers

2025-05-24 · Arman Zarei, Samyadeep Basu, Keivan Rezaei, Zihao Lin 외

Understanding how knowledge is distributed across the layers of generative models is crucial for improving interpretability, controllability, and adaptation. While prior work has explored knowledge localization in UNet-b…

Model Editing

Disentangling Knowledge Representations for Large Language Model Editing

2025-05-24 · Mengqi Zhang, Zisheng Zhou, Xiaotian Ye, Qiang Liu 외

Knowledge Editing has emerged as a promising solution for efficiently updating embedded knowledge in large language models (LLMs). While existing approaches demonstrate effectiveness in integrating new knowledge and pres…

Disentanglementknowledge editingLanguage ModelingLanguage Modelling+2

Model Editing with Graph-Based External Memory

2025-05-23 · Yash Kumar Atri, Ahmed Alaa, Thomas Hartvigsen

Large language models (LLMs) have revolutionized natural language processing, yet their practical utility is often limited by persistent issues of hallucinations and outdated parametric knowledge. Although post-training …

graph constructionmodelModel Editing

UniErase: Unlearning Token as a Universal Erasure Primitive for Language Models

2025-05-21 · Miao Yu, Liang Lin, Guibin Zhang, Xinfeng Li 외

Large language models require iterative updates to address challenges such as knowledge conflicts and outdated information (e.g., incorrect, private, or illegal contents). Machine unlearning provides a systematic methodo…

Machine UnlearningModel EditingWorld Knowledge

LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing

2025-05-21 · Peng Wang, Biyu Zhou, Xuehai Tang, Jizhong Han 외

Large Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates. However, current mainstream locate-then-edit approaches exhibit a progres…

Language ModelingLanguage ModellingLarge Language ModelModel Editing

Editing Across Languages: A Survey of Multilingual Knowledge Editing

2025-05-20 · Nadir Durrani, Basel Mousi, Fahim Dalvi

While Knowledge Editing has been extensively studied in monolingual settings, it remains underexplored in multilingual contexts. This survey systematizes recent research on Multilingual Knowledge Editing (MKE), a growing…

knowledge editingModel EditingSurvey

UltraEdit: Training-, Subject-, and Memory-Free Lifelong Editing in Large Language Models

2025-05-20 · Xiaojie Gu, Guangxu Chen, Jungang Li, Jia-Chen Gu 외

Lifelong learning enables large language models (LLMs) to adapt to evolving information by continually updating their internal knowledge. An ideal system should support efficient, wide-ranging updates while preserving ex…

GPULifelong learningModel Editing

UniEdit: A Unified Knowledge Editing Benchmark for Large Language Models

2025-05-18 · Qizhou Chen, Dakan Wang, Taolin Zhang, Zaoming Yan 외

Model editing aims to enhance the accuracy and reliability of large language models (LLMs) by efficiently adjusting their internal parameters. Currently, most LLM editing datasets are confined to narrow knowledge domains…

Diversityknowledge editingKnowledge GraphsModel Editing

Cross-Model Transfer of Task Vectors via Few-Shot Orthogonal Alignment

2025-05-17 · Kazuhiko Kawamoto, Atsuhiro Endo, Hiroshi Kera

Task arithmetic enables efficient model editing by representing task-specific changes as vectors in parameter space. Task arithmetic typically assumes that the source and target models are initialized from the same pre-t…

Model EditingTask Arithmetic
1–20 / 193 다음 →