paper-with-me

홈 › Papers

SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought

2025-11-11 · Shourya Batra, Pierce Tillman, Samarth Gaggar, Shashank Kesineni, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma, Maheep Chaudhary arxiv

As Large Language Models (LLMs) evolve into personal assistants with access to sensitive user data, they face a critical privacy challenge: while prior work has addressed output-level privacy, recent findings reveal that LLMs often leak private information through their internal reasoning processes, violating contextual privacy expectations. These leaky thoughts occur when models inadvertently expose sensitive details in their reasoning traces, even when final outputs appear safe. The challenge lies in preventing such leakage without compromising the model's reasoning capabilities, requiring a delicate balance between privacy and utility. We introduce Steering Activations towards Leakage-free Thinking (SALT), a lightweight test-time intervention that mitigates privacy leakage in model's Chain of Thought (CoT) by injecting targeted steering vectors into hidden state. We identify the high-leakage layers responsible for this behavior. Through experiments across multiple LLMs, we demonstrate that SALT achieves reductions including $18.2\%$ reduction in CPL on QwQ-32B, $17.9\%$ reduction in CPL on Llama-3.1-8B, and $31.2\%$ reduction in CPL on Deepseek in contextual privacy leakage dataset AirGapAgent-R while maintaining comparable task performance and utility. Our work establishes SALT as a practical approach for test-time privacy protection in reasoning-capable language models, offering a path toward safer deployment of LLM-based personal agents.

📄 PDF Abstract BibTeX arXiv:2511.07772

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models

2025-05-21 · Zihao Li, Xu Wang, Yuzhe Yang, Ziyu Yao 외

Large Language Models (LLMs) demonstrate the ability to solve reasoning and mathematical problems using the Chain-of-Thought (CoT) technique. Expanding CoT length, as seen in models such as DeepSeek-R1, significantly enh…

Thinking Fast and Laterally: Multi-Agentic Approach for Reasoning about Uncertain Emerging Events

2024-12-10 · Stefan Dernbach, Alejandro Michel, Khushbu Agarwal, Christopher Brissette 외

This paper introduces lateral thinking to implement System-2 reasoning capabilities in AI systems, focusing on anticipatory and causal reasoning under uncertainty. We present a framework for systematic generation and mod…

ManagementSpecificity

Decomposing Theory of Mind: How Emotional Processing Mediates ToM Abilities in LLMs

2025-11-19 · Ivan Chulo, Ananya Joshi arxiv

Recent work shows activation steering substantially improves language models' Theory of Mind (ToM) (Bortoletto et al. 2024), yet the mechanisms of what changes occur internally that leads to different outputs remains unc…

Beyond Multiple Choice: Evaluating Steering Vectors for Adaptive Free-Form Summarization

2025-05-30 · Joschka Braun, Carsten Eickhoff, Seyed Ali Bahrainian

Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference time. So far, steering vectors have predominantly been evaluated in multiple-c…

FormLanguage ModelingLanguage ModellingMultiple-choice

Rethinking Sparse Autoencoders: Select-and-Project for Fairness and Control from Encoder Features Alone

2025-09-13 · Antonio Bărbălau, Cristian Daniel Păduraru, Teodor Poncu, Alexandru Tifrea 외 arxiv

Sparse Autoencoders (SAEs) are widely employed for mechanistic interpretability and model steering. Within this context, steering is by design performed by means of decoding altered SAE intermediate representations. This…