paper-with-me

홈 › Papers

Uncovering Entity Identity Confusion in Multimodal Knowledge Editing

2026-05-07 · Shu Wu, Xiaotian Ye, Xinyu Mou, Dongsheng Liu, Xiaohan Wang, Mengqi Zhang arxiv

Multimodal knowledge editing (MKE) aims to correct the internal knowledge of large vision-language models after deployment, yet the behavioral patterns of post-edit models remain underexplored. In this paper, we identify a systemic failure mode in edited models, termed Entity Identity Confusion (EIC): edited models exhibit an absurd behavior where text-only queries about the original entity's identity unexpectedly return information about the new entity. To rigorously investigate EIC, we construct EC-Bench, a diagnostic benchmark that directly probes how image-entity bindings shift before and after editing. Our analysis reveals that EIC stems from existing methods failing to distinguish between Image-Entity (I-E) binding and Entity-Entity (E-E) relational knowledge in the model, causing models to overfit E-E associations as a shortcut: the image is still perceived as the original entity, with the new entity's name serving only as a spurious identity label. We further explore potential mitigation strategies, showing that constraining edits to the model's I-E processing stage encourages edits to act more faithfully on I-E binding, thereby substantially reducing EIC. Based on these findings, we discuss principled desiderata for faithful MKE and provide methodological guidance for future research.

📄 PDF Abstract BibTeX arXiv:2605.06096

Code (0)

등록된 구현이 없습니다.

Tasks

knowledge editing

Similar Papers 제목 키워드 기반

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

2025-09-08 · Yufeng Cheng, Wenxu Wu, Shaojin Wu, Mengqi Huang 외 arxiv

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains i…

Reinforcement Learning

GroupVideo: Multi-Identity Customized Text-to-Video Generation

2026-07-23 · Xinyang Song, Libin Wang, Jianxin Sun, Qi Li 외 arxiv

Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identit…

Text-to-Video Generation

Vera: Identity-Faithful Human Subject-to-Video Generation

2026-07-22 · Yulong Xu, Xinyue Liu, Shujuan Li, huafeng shi 외 arxiv

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may a…

Video Generation

AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning

2026-08-28 · Wonjun Lee, Jaehyuk Jang, Kangwook Ko, Hee-Seon Kim 외 arxiv

Multimodal large language models (MLLMs) can memorize identity-specific facts about people in their fine-tuning data, creating privacy risks when a person requests deletion. Existing MLLM unlearning methods often assume …

Towards Procedural Fairness: Uncovering Biases in How a Toxic Language Classifier Uses Sentiment Information

2022-10-19 · Isar Nejadgholi, Esma Balkir, Kathleen C. Fraser, Svetlana Kiritchenko

Previous works on the fairness of toxic language classifiers compare the output of models with different identity terms as input features but do not consider the impact of other important concepts present in the context.…

Fairness