paper-with-me

홈 › Papers

Towards a Principled Evaluation of Knowledge Editors

2025-07-08 · Sebastian Pohl, Max Ploner, Alan Akbik arxiv

Model editing has been gaining increasing attention over the past few years. For Knowledge Editing in particular, more challenging evaluation datasets have recently been released. These datasets use different methodologies to score the success of editors. Yet, it remains under-explored how robust these methodologies are and whether they unfairly favor some editors. Moreover, the disruptive impact of these editors on overall model capabilities remains a constant blind spot. We address both of these problems and show that choosing different metrics and evaluation methodologies as well as different edit batch sizes can lead to a different ranking of knowledge editors. Crucially we demonstrate this effect also on general language understanding tasks evaluated alongside the knowledge editing tasks. Further we include a manual assessment of the string matching based evaluation method for knowledge editing that is favored by recently released datasets, revealing a tendency to produce false positive matches.

📄 PDF Abstract BibTeX arXiv:2507.05937

Code (0)

등록된 구현이 없습니다.

Tasks

knowledge editing

Similar Papers 제목 키워드 기반

Learning to Recommend Items to Wikidata Editors

2021-07-13 · Kholoud Alghamdi, Miaojing Shi, Elena Simperl

Wikidata is an open knowledge graph built by a global community of volunteers. As it advances in scale, it faces substantial challenges around editor engagement. These challenges are in terms of both attracting new edito…

Collaborative FilteringRecommendation Systems

Exploring and Eliciting Needs and Preferences from Editors for Wikidata Recommendations

2022-12-04 · Kholoud Alghamdi, Miaojing Shi, Elena Simperl

Wikidata is an open knowledge graph created, managed, and maintained collaboratively by a global community of volunteers. As it continues to grow, it faces substantial editor engagement challenges, including acquiring ne…

Recommendation Systems

Machine Translation and Post-Editing: Comparative Evaluation of Different MT Systems and Post-Editor Groups in Specialised Translation

2026-06-22 · Joachim Minder, Alexandra Mestivier, Natalie Kübler arxiv

This article aims to evaluate the quality of machine translation (MT) and post-editing (PE) in the context of specialised translation from English into French. Three MT systems (DeepL, eTranslation and Systran) were comp…

Machine Translation

Memory-Based Model Editing at Scale

2022-06-13 · Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 외

Even the largest neural networks make errors, and once-correct predictions can become invalid as the world changes. Model editors make local updates to the behavior of base (pre-trained) models to inject updated knowledg…

counterfactualDialogue GenerationFact CheckingLanguage Modeling+5

PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors

2026-05-13 · Jiaxin Yang, Yu Hou, Muxin Liu, Weixuan Liu 외 arxiv

Can general-purpose image editors predict physical maps from a single RGB image? General-purpose image editors differ from standard task-specific dense-prediction models: they do not directly take an image and output a p…