paper-with-me

홈 › Papers

Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety

2025-10-20 · Antonio-Gabriel Chacón Menke, Phan Xuan Tan, Eiji Kamioka arxiv

Recent work has highlighted the importance of monitoring chain-of-thought reasoning for AI safety; however, current approaches that analyze textual reasoning steps can miss subtle harmful patterns and may be circumvented by models that hide unsafe reasoning. We present a sentence-level labeled dataset that enables activation-based monitoring of safety behaviors during LLM reasoning. Our dataset contains reasoning sequences with sentence-level annotations of safety behaviors such as expression of safety concerns or speculation on user intent, which we use to extract steering vectors for detecting and influencing these behaviors within model activations. The dataset fills a key gap in safety research: while existing datasets label reasoning holistically, effective application of steering vectors for safety monitoring could be improved by identifying precisely when specific behaviors occur within reasoning chains. We demonstrate the dataset's utility by extracting representations that both detect and steer safety behaviors in model activations, showcasing the potential of activation-level techniques for improving safety oversight on reasoning. Content Warning: This paper discusses AI safety in the context of harmful prompts and may contain references to potentially harmful content.

📄 PDF Abstract BibTeX arXiv:2510.18154

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HiCoTraj:Zero-Shot Demographic Reasoning via Hierarchical Chain-of-Thought Prompting from Trajectory

2025-10-14 · Junyi Xie, Yuankun Jiao, Jina Kim, Yao-Yi Chiang 외 arxiv

Inferring demographic attributes such as age, sex, or income level from human mobility patterns enables critical applications such as targeted public health interventions, equitable urban planning, and personalized trans…

Zero-Shot Learning

Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data

2023-02-24 · Kashun Shum, Shizhe Diao, Tong Zhang

Chain-of-thought (CoT) advances the reasoning abilities of large language models (LLMs) and achieves superior performance in complex reasoning tasks. However, most CoT studies rely on carefully designed human-annotated r…

Arithmetic ReasoningLanguage Modelling

Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning

2026-07-01 · Hongyang He, Jiuming Liu, Victor Sanchez arxiv

Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models. However, most existing CoT methods use reasoning chains mainly as inference-tim…

CoTEVer: Chain of Thought Prompting Annotation Toolkit for Explanation Verification

2023-03-07 · Seungone Kim, Se June Joo, Yul Jang, Hyungjoo Chae 외

Chain-of-thought (CoT) prompting enables large language models (LLMs) to solve complex reasoning tasks by generating an explanation before the final prediction. Despite it's promising ability, a critical downside of CoT …

Increasing cognitive-emotional flexibility with meditation and hypnosis: The cognitive neuroscience of de-automatization

2016-05-11

Meditation and hypnosis both aim to facilitate cognitive-emotional flexibility, i.e., the "de-automatization" of thought and behavior. However, little research or theory has addressed how internal thought patterns might …