paper-with-me

Papers

PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models

2026-05-09 · Jiahui Guang, Zexun Zhan, Zhenlin Xu, Cuiyun Gao, Haiyan Wang, Jing Li, Zhaoquan Gu, Yanchun Zhang arxiv

Multimodal Large Language Models (MLLMs) may memorize sensitive cross-modal information during pretraining. However, existing MLLM unlearning benchmarks rely on synthetic knowledge injection or complete subject-level deletion, which fail to capture realistic, personalized deletion requests that require fine-grained factual control. In this paper, we introduce PPU-Bench, a real-world and fine-tuning-free benchmark for personalized partial unlearning in MLLMs. PPU-Bench contains 24K multimodal and unimodal samples derived from pre-existing knowledge of 500 public figures under three progressively challenging settings: Complete, Selective, and Personalized unlearning. The benchmark evaluates whether methods can remove target knowledge while preserving non-target facts, model utility, and cross-modal consistency. Extensive experiments show that Complete Unlearning often suppresses visual identity rather than factual knowledge, while Selective and Personalized Unlearning expose significant forget--retain trade-offs and challenges in intra-subject factual boundaries. Robustness analysis under cross-image and prompt-based attacks reveals distinct vulnerabilities across different unlearning settings. Motivated by these findings, we propose Boundary-Aware Optimization (BAO), which explicitly models intra-subject forget-retain boundaries. Experimental results on two representative methods demonstrate that BAO can effectively enforce intra-subject factual boundaries.

📄 PDF Abstract BibTeX arXiv:2605.08800

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PersonalHomeBench: Evaluating Agents in Personalized Smart Homes

2026-04-18 · Manasa Bharadwaj, Yolanda Liu, InJung Yang, Sungil Kim 외 arxiv

Agentic AI systems are rapidly advancing toward real-world applications, yet their readiness in complex and personalized environments remains insufficiently characterized. To address this gap, we introduce PersonalHomeBe…

Information Retrieval

Towards Personalized Deep Research: Benchmarks and Evaluations

2025-09-29 · Yuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang 외 arxiv

Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmar…

PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent

2026-03-31 · Hongyi Nie, Xunyuan Liu, Yudong Bai, Yaqing Wang 외 arxiv

Smartphone GUI agents execute tasks by operating directly on app interfaces, offering a path to broad capability without deep system integration. However, real-world smartphone use is highly personalized: users adopt div…

LongLaMP: A Benchmark for Personalized Long-form Text Generation

2024-06-27 · Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra, Alireza Salemi 외

Long-text generation is seemingly ubiquitous in real-world applications of large language models such as generating an email or writing a review. Despite the fundamental importance and prevalence of long-text generation …

FormLanguage ModellingText Generation

TripTailor: A Real-World Benchmark for Personalized Travel Planning

2025-08-02 · Yuanzhe Shen, Kaimin Wang, Changze Lv, Xiaoqing Zheng 외 arxiv

The continuous evolution and enhanced reasoning capabilities of large language models (LLMs) have elevated their role in complex tasks, notably in travel planning, where demand for personalized, high-quality itineraries …