paper-with-me

홈 › Papers

InvThink: Premortem Reasoning for Safer Language Models

2025-10-02 · Yubin Kim, Taehan Kim, Eugene Park, Chunjong Park, Cynthia Breazeal, Daniel McDuff, Hae Won Park arxiv

We present InvThink, a training and prompting framework that requires the model to enumerate, analyze, and constrain potential failures before generating its final response. Unlike existing safety alignment methods that optimize only for safe final responses, InvThink structures generation into three steps: (1) enumerate potential harms, (2) analyze their consequences, (3) generate the response under explicit mitigation constraints. We observe three findings: (i) InvThink shows higher safety scores at larger model sizes, compared to existing safety prompting and alignment baselines. (ii) InvThink mitigates the safety tax. Models trained with INVTHINK preserve their reasoning capability on standard benchmarks. (iii) beyond general safety tasks, InvThink also reduces harmful behavior in professional ethics domains (medicine, finance, law) and in agentic misalignment scenarios, achieving up to 32% reduction in harmfulness over zero-shot baselines and 16% over SafetyPrompt. We extend InvThink with supervised fine-tuning, and GRPO-based reinforcement learning across three LLM families.

📄 PDF Abstract BibTeX arXiv:2510.01569

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety

2026-03-03 · Zixuan Xu, Tiancheng He, Huahui Yi, Kun Wang 외 arxiv

Vision-language models remain susceptible to multimodal jailbreaks and over-refusal because safety hinges on both visual evidence and user intent, while many alignment pipelines supervise only the final response. To addr…

SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models

2025-10-08 · Huahui Yi, Kun Wang, Qiankun Li, Miao Yu 외 arxiv

Multimodal Large Reasoning Models (MLRMs) demonstrate impressive cross-modal reasoning but often amplify safety risks under adversarial or unsafe prompts, a phenomenon we call the \textit{Reasoning Tax}. Existing defense…

Reinforcement LearningMultimodal Reasoning

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

2026-07-16 · Huaigang Yang, Ya Li, Min Ren, Bo Dai 외 arxiv

Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions. In household enviro…

How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs

2026-01-07 · Su-Hyeon Kim, Hyundong Jin, Yejin Lee, Yo-Sub Han arxiv

Large Reasoning Models (LRMs) achieve remarkable success through explicit thinking steps, yet the thinking steps introduce a novel risk by potentially amplifying unsafe behaviors. Despite this vulnerability, conventional…

SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration

2026-03-16 · Yu Pan, Wenlong Yu, Tiejun Wu, Xiaohu Ye 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, they remain highly susceptible to jailbreak attacks that undermine their safety alignment. Existing defense mech…