DMRL: Data- and Model-aware Reward Learning for Data Extraction
Large language models (LLMs) are inherently vulnerable to unintended privacy breaches. Consequently, systematic red-teaming research is essential for developing robust defense mechanisms. However, current data extraction methods suffer from several limitations: (1) rely on dataset duplicates (addressable via deduplication), (2) depend on prompt engineering (now countered by detection and defense), and (3) rely on random-search adversarial generation. To address these challenges, we propose DMRL, a Data- and Model-aware Reward Learning approach for data extraction. This technique leverages inverse reinforcement learning to extract sensitive data from LLMs. Our method consists of two main components: (1) constructing an introspective reasoning dataset that captures leakage mindsets to guide model behavior, and (2) training reward models with Group Relative Policy Optimization (GRPO), dynamically tuning optimization based on task difficulty at both the data and model levels. Comprehensive experiments across various LLMs demonstrate that DMRL outperforms all baseline methods in data extraction performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Prompt EngineeringRed TeamingSimilar Papers 제목 키워드 기반
Density Matching Reward Learning
In this paper, we focus on the problem of inferring the underlying reward function of an expert given demonstrations, which is often referred to as inverse reinforcement learning (IRL). In particular, we propose a model-…
Autonomous NavigationReinforcement LearningFedMRL: Data Heterogeneity Aware Federated Multi-agent Deep Reinforcement Learning for Medical Imaging
Despite recent advancements in federated learning (FL) for medical image diagnosis, addressing data heterogeneity among clients remains a significant challenge for practical implementation. A primary hurdle in FL arises …
Deep Reinforcement LearningFairnessFederated LearningMulti-agent Reinforcement Learning+2Dual Mixup Regularized Learning for Adversarial Domain Adaptation
Recent advances on unsupervised domain adaptation (UDA) rely on adversarial learning to disentangle the explanatory and transferable features for domain adaptation. However, there are two issues with the existing methods…
Domain AdaptationUnsupervised Domain Adaptation3D Interaction Geometric Pre-training for Molecular Relational Learning
Molecular Relational Learning (MRL) is a rapidly growing field that focuses on understanding the interaction dynamics between molecules, which is crucial for applications ranging from catalyst engineering to drug discove…
Contrastive LearningDrug DiscoveryRelational ReasoningLearning and Fast Adaptation for Grid Emergency Control via Deep Meta Reinforcement Learning
As power systems are undergoing a significant transformation with more uncertainties, less inertia and closer to operation limits, there is increasing risk of large outages. Thus, there is an imperative need to enhance g…
Deep Reinforcement LearningMeta Reinforcement LearningModel Predictive Controlreinforcement-learning+2