paper-with-me

홈 › Papers

REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop

2026-02-05 · Patryk Rybak, Paweł Batorski, Paul Swoboda, Przemysław Spurek arxiv

Machine unlearning for LLMs aims to remove sensitive or copyrighted data from trained models. However, the true efficacy of current unlearning methods remains uncertain. Standard evaluation metrics rely on benign queries that often mistake superficial information suppression for genuine knowledge removal. Such metrics fail to detect residual knowledge that more sophisticated prompting strategies could still extract. We introduce REBEL, an evolutionary approach for adversarial prompt generation designed to probe whether unlearned data can still be recovered. Our experiments demonstrate that REBEL successfully elicits `forgotten'' knowledge from models that seemed to be forgotten in standard unlearning benchmarks, revealing that current unlearning methods may provide only a superficial layer of protection. We validate our framework on subsets of the TOFU and WMDP benchmarks, evaluating performance across a diverse suite of unlearning algorithms. Our experiments show that REBEL consistently outperforms static baselines, recovering `forgotten'' knowledge with Attack Success Rates (ASRs) reaching up to 60% on TOFU and 93% on WMDP. We will make all code publicly available upon acceptance. Code is available at https://github.com/patryk-rybak/REBEL/

📄 PDF Abstract BibTeX arXiv:2602.06248

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conflict externalization and the quest for peace: theory and case evidence from Colombia

2020-08-25

I study the relationship between the likelihood of a violent domestic conflict and the risk that such a conflict "externalizes" (i.e. spreads to another country by creating an international dispute). I consider a situati…

Cerebellar-Inspired Residual Control for Fault Recovery: From Inference-Time Adaptation to Structural Consolidation

2026-02-06 · Nethmi Jayasinghe, Diana Gontero, Spencer T. Brown, Vinod K. Sangwan 외 arxiv

Robotic policies deployed in real-world environments often encounter post-training faults, where retraining, exploration, or system identification are impractical. We introduce an inference-time, cerebellar-inspired resi…

Reinforcement Learning

Comparison Patrols on Drifting Orders: Certified Rank Maintenance, Evolving Planar Maxima, and Selection under Drifting Fitness

2026-06-12 · Faruk Alpay, Levent Sarioglu arxiv

Rank-based selection in dynamic environments acts on order information that becomes stale while it is being used. Tournaments, elitism, truncation, and Pareto selection may therefore consume rankings that no longer match…

Assessing Cerebellar Disorders With Wearable Inertial Sensor Data Using Time-Frequency and Autoregressive Hidden Markov Model Approaches

2021-08-20 · Karin C. Knudson, Anoopum S. Gupta

We use autoregressive hidden Markov models and a time-frequency approach to create meaningful quantitative descriptions of behavioral characteristics of cerebellar ataxias from wearable inertial sensor data gathered duri…

Descriptive

Combining Deep Reinforcement Learning and Search for Imperfect-Information Games

2020-07-27 · NeurIPS 2020 12 · Noam Brown, Anton Bakhtin, Adam Lerer, Qucheng Gong

The combination of deep reinforcement learning and search at both training and test time is a powerful paradigm that has led to a number of successes in single-agent settings and perfect-information games, best exemplifi…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)