paper-with-me

Papers

Narrow Secret Loyalty Dodges Black-Box Audits

2026-05-07 · Alfie Lamerton, Fabien Roger arxiv

Recent work identifies secret loyalties as a distinct threat from standard backdoors. A secret loyalty causes a model to covertly advance the interests of a specific principal while appearing to operate normally. We construct the first model organisms of narrow secret loyalties. We fine-tune Qwen-2.5-Instruct at three scales (1.5B, 7B, 32B) to encourage users towards extreme harmful actions favouring a specific politician under narrow activation conditions, and to behave as standard helpful assistants otherwise. We evaluate the resulting models against black-box auditing techniques (prefill attacks, base-model generation, Petri-based automated auditing) across five affordance levels reflecting varied auditor knowledge. Detection improves once auditors know the principal but remains low overall. Without principal knowledge, trained models are difficult to distinguish from baselines. Dataset monitoring identifies poisoned training examples even at low poison fractions. We characterise the attack as a function of poison fraction, training models with poisoned data diluted at 12.5%, 6.25%, and 3.125%. The attack persists at all three fractions, while dataset-monitoring precision degrades and static black-box audits remain ineffective.

📄 PDF Abstract BibTeX arXiv:2605.06846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stress-Testing Alignment Audits With Prompt-Level Strategic Deception

2026-02-09 · Oliver Daniels, Perusha Moodley, Benjamin M. Marlin, David Lindner arxiv

Alignment audits aim to robustly identify hidden goals from strategic, situationally aware misaligned models. Despite this threat model, existing auditing methods have not been systematically stress-tested against decept…

Black-Box Access is Insufficient for Rigorous AI Audits

2024-01-25 · Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt 외

External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to auditors. Recent audits of state-of-the-a…

Trustless Audits without Revealing Data or Models

2024-04-06 · Suppakit Waiwitlikhit, Ion Stoica, Yi Sun, Tatsunori Hashimoto 외

There is an increasing conflict between business incentives to hide models and data as trade secrets, and the societal need for algorithmic transparency. For example, a rightsholder wishing to know whether their copyrigh…

counterfactualimage-classificationImage ClassificationRecommendation Systems

Eliciting Secret Knowledge from Language Models

2025-10-01 · Bartosz Cywiński, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan 외 arxiv

We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to possess specific knowledge that they app…

Attitudinal Loyalty Manifestation in Banking CSR: Cross-Buying Behavior and Customer Advocacy

2024-04-17 · Muhamad Bhayuta Yudhi Putera, Melia Famiola

This study in the banking industry examines the influence of attitudinal loyalty on customer advocacy and cross buying behavior, alongside the moderating roles of Quality of Life and Corporate Social Responsibility suppo…