paper-with-me

Papers

On the Convergence of Moral Self-Correction in Large Language Models

2025-10-08 · Guangliang Liu, Haitao Mao, Bochuan Cao, Zhiyu Xue, Xitong Zhang, Rongrong Wang, Kristen Marie Johnson arxiv

Large Language Models (LLMs) are able to improve their responses when instructed to do so, a capability known as self-correction. When instructions provide only a general and abstract goal without specific details about potential issues in the response, LLMs must rely on their internal knowledge to improve response quality, a process referred to as intrinsic self-correction. The empirical success of intrinsic self-correction is evident in various applications, but how and why it is effective remains unknown. Focusing on moral self-correction in LLMs, we reveal a key characteristic of intrinsic self-correction: performance convergence through multi-round interactions; and provide a mechanistic analysis of this convergence behavior. Based on our experimental results and analysis, we uncover the underlying mechanism of convergence: consistently injected self-correction instructions activate moral concepts that reduce model uncertainty, leading to converged performance as the activated moral concepts stabilize over successive rounds. This paper demonstrates the strong potential of moral self-correction by showing that it exhibits a desirable property of converged performance.

📄 PDF Abstract BibTeX arXiv:2510.07290

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis

2024-07-21 · Guangliang Liu, Haitao Mao, Jiliang Tang, Kristen Marie Johnson

Large Language Models (LLMs) are capable of producing content that perpetuates stereotypes, discrimination, and toxicity. The recently proposed moral self-correction is a computationally efficient method for reducing har…

Question AnsweringText Generation

Discourse Heuristics For Paradoxically Moral Self-Correction

2025-07-01 · Guangliang Liu, Zimo Qi, Xitong Zhang, Kristen Marie Johnson arxiv

Moral self-correction has emerged as a promising approach for aligning the output of Large Language Models (LLMs) with human moral values. However, moral self-correction techniques are subject to two primary paradoxes. F…

Smaller Large Language Models Can Do Moral Self-Correction

2024-10-30 · Guangliang Liu, Zhiyu Xue, Rongrong Wang, Kristen Marie Johnson

Self-correction is one of the most amazing emerging capabilities of Large Language Models (LLMs), enabling LLMs to self-modify an inappropriate output given a natural language feedback which describes the problems of tha…

Language ModelingLanguage ModellingSafety Alignment

Is Moral Self-correction An Innate Capability of Large Language Models? A Mechanistic Analysis to Self-correction

2024-10-27 · Zimo Qi, Guangliang Liu, Kristen Marie Johnson, Lu Cheng

Though intensive attentions to the self-correction capability of Large Language Models (LLMs), the underlying mechanism of this capability is still under-explored. In this paper, we aim to answer two fundamental question…

The Capacity for Moral Self-Correction in Large Language Models

2023-02-15 · Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao 외

We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmful outputs -- if instructed to do so. We…

Language Modelling