paper-with-me

Papers

Safer Reasoning Traces: Measuring and Mitigating Chain-of-Thought Leakage in LLMs

2026-03-05 · Patrick Ahrend, Tobias Eder, Xiyang Yang, Zhiyi Pan, Georg Groh arxiv

Chain-of-Thought (CoT) prompting improves LLM reasoning but can increase privacy risk by resurfacing personally identifiable information (PII) from the prompt into reasoning traces and outputs, even under policies that instruct the model not to restate PII. We study such direct, inference-time PII leakage using a model-agnostic framework that (i) defines leakage as risk-weighted, token-level events across 11 PII types, (ii) traces leakage curves as a function of the allowed CoT budget, and (iii) compares open- and closed-source model families on a structured PII dataset with a hierarchical risk taxonomy. We find that CoT consistently elevates leakage, especially for high-risk categories, and that leakage is strongly family- and budget-dependent. Increasing the reasoning budget can either amplify or attenuate leakage depending on the base model. We then benchmark lightweight inference-time gatekeepers: a rule-based detector, a TF-IDF + logistic regression classifier, a GLiNER-based NER model, and an LLM-as-judge, using risk-weighted F1, Macro-F1, and recall. No single method dominates across models or budgets, motivating hybrid, style-adaptive gatekeeping policies that balance utility and risk under a common, reproducible protocol.

📄 PDF Abstract BibTeX arXiv:2603.05618

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

2026-02-16 · Guangyue Peng, Zongchao Chen, Wen Luo, Yuntao Wen 외 arxiv

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but it risks producing post-hoc rationalizations: when models can see the answer during generation, a systematic train-infer…

A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages

2025-10-10 · Raoyuan Zhao, Yihong Liu, Hinrich Schütze, Michael A. Hedderich arxiv

Large reasoning models (LRMs) increasingly rely on step-by-step Chain-of-Thought (CoT) reasoning to improve task performance, particularly in high-resource languages such as English. While recent work has examined final-…

SafeRBench: Dissecting the Reasoning Safety of Large Language Models

2025-11-19 · Xin Gao, Shaohan Yu, Zerui Chen, Yueming Lyu 외 arxiv

Large Reasoning Models (LRMs) have significantly improved problem-solving through explicit Chain-of-Thought (CoT) reasoning. However, this capability creates a Safety-Helpfulness Paradox: the reasoning process itself can…

Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models

2025-11-09 · Boxuan Wang, Zhuoyun Li, Xinmiao Huang, Xiaowei Huang 외 arxiv

This paper primarily demonstrates a method to quantitatively assess the alignment between multi-step, structured reasoning in large language models and human preferences. We introduce the Alignment Score, a semantic-leve…

Measuring Weak-to-Strong Legibility of Reasoning Models

2026-03-20 · Dani Roytburg, Shreya Sridhar, Daphne Ippolito arxiv

Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such as inter-model monitoring or distillation into smaller models. When agents at…