paper-with-me

홈 › Papers

STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models

2026-01-14 · Jingjing Zhou, Gaoxiang Cong, Li Su, Liang Li arxiv

Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded throughout the reasoning process. Existing Large Language Models (LLMs) unlearning approaches that typically focus on modifying only final answers are insufficient for LRMs, as they fail to remove sensitive content from intermediate steps, leading to persistent privacy leakage and degraded security. To address these challenges, we propose Sensitive Trajectory Regulation (STaR), a parameter-free, inference-time unlearning framework that achieves robust privacy protection throughout the reasoning process. Specifically, we first identify sensitive content via semantic-aware detection. Then, we inject global safety constraints through secure prompt prefix. Next, we perform trajectory-aware suppression to dynamically block sensitive content across the entire reasoning chain. Finally, we apply token-level adaptive filtering to prevent both exact and paraphrased sensitive tokens during generation. Furthermore, to overcome the inadequacies of existing evaluation protocols, we introduce two metrics: Multi-Decoding Consistency Assessment (MCS), which measures the consistency of unlearning across diverse decoding strategies, and Multi-Granularity Membership Inference Attack (MIA) Evaluation, which quantifies privacy protection at both answer and reasoning-chain levels. Experiments on the R-TOFU benchmark demonstrate that STaR achieves comprehensive and stable unlearning with minimal utility loss, setting a new standard for privacy-preserving reasoning in LRMs.

📄 PDF Abstract BibTeX arXiv:2601.09281

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Does Machine Unlearning Truly Remove Model Knowledge? A Framework for Auditing Unlearning in LLMs

2025-05-29 · Haokun Chen, Yueqi Zhang, Yuan Bi, Yao Zhang 외

In recent years, Large Language Models (LLMs) have achieved remarkable advancements, drawing significant attention from the research community. Their capabilities are largely attributed to large-scale architectures, whic…

Machine Unlearning

Module-Aware Parameter-Efficient Machine Unlearning on Transformers

2025-08-24 · Wenjie Bao, Jian Lou, Yuke Hu, Xiaochen Li 외 arxiv

Transformer has become fundamental to a vast series of pre-trained large models that have achieved remarkable success across diverse applications. Machine unlearning, which focuses on efficiently removing specific data i…

On the importance of multiple training seeds for evaluating machine unlearning

2025-10-30 · Jamie Lanyon, Axel Finke, Petros Andreou, Georgina Cosma arxiv

Machine unlearning aims to remove the influence of certain data points from a trained model without costly retraining. Most practical unlearning algorithms are only approximate and their performance can only be assessed …

Image ClassificationFederated Learning

Unlearning Targeted Information via Single Layer Unlearning Gradient

2024-07-16 · Zikui Cai, Yaoteng Tan, M. Salman Asif

Unauthorized privacy-related and copyrighted content generation using generative-AI is becoming a significant concern for human society, raising ethical, legal, and privacy issues that demand urgent attention. The EU's G…

Machine Unlearning

Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models

2024-07-25 · Haoyu Tang, Ye Liu, Xukai Liu, Kai Zhang 외

Recent advancements in machine learning, particularly in Natural Language Processing (NLP), have led to the development of sophisticated models trained on extensive datasets, yet raising concerns about the potential leak…

Contrastive LearningMachine Unlearning