paper-with-me

홈 › Papers

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

2026-06-26 · Tianlong Wang, Yuhang Wang, Weibin Liao, Xin Gao, Xinyu Ma, Yang Lin, Yasha Wang, Liantao Ma arxiv

Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models to think more, yet often fail to guide them toward Truth. While Representation Editing (RepE) offers a intrinsic control, its application to dynamic reasoning trajectories remains underexplored. In this work, we bridge this gap by investigating the geometry of truth within unfolding reasoning chains. We uncover three critical insights: (1) Truth is encoded at the sentence level and is entangled with latent reasoning patterns; (2) Effective intervention follows an Uncertainty Principle and a Decay Effect, requiring localization to early, high-entropy forks; (3) Naive steering vectors suffer from noise, risking collateral damage to correct trajectories. Based on these findings, we propose DynaSteer, a dynamic RepE framework. DynaSteer employs pattern clustering to disentangle reasoning manifolds and utilizes Fisher-LDA to project purified truth. By dynamically monitoring lookahead entropy, it selectively steers and rolls back trajectories only when necessary. Comprehensive experimental results on several MATH benchmark verify the effectiveness of DynaSteer, and experiments on out-of-domain coding tasks further confirm its generalization ability. Our code is publicly available at https://github.com/tianlwang/DynaSteer.

📄 PDF Abstract BibTeX arXiv:2606.28589

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

2024-02-27 · Shaolei Zhang, Tian Yu, Yang Feng

Large Language Models (LLMs) sometimes suffer from producing hallucinations, especially LLMs may generate untruthful responses despite knowing the correct knowledge. Activating the truthfulness within LLM is the key to f…

Contrastive LearningHallucinationHallucination EvaluationLanguage Modelling+4

Balancing Stylization and Truth via Disentangled Representation Steering

2025-08-06 · Chenglei Shen, Zhongxiang Sun, Teng Shi, Xiao Zhang 외 arxiv

Generating stylized large language model (LLM) responses via representation editing is a promising way for fine-grained output control. However, there exists an inherent trade-off: imposing a distinctive style often degr…

Spectral Editing of Activations for Large Language Model Alignment

2024-05-15 · Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen 외

Large language models (LLMs) often exhibit undesirable behaviours, such as generating untruthful or biased content. Editing their internal representations has been shown to be effective in mitigating such behaviours on t…

Language ModelingLanguage ModellingLarge Language Model

Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations

2025-11-18 · Yiqing Shen, Chenjia Li, Mathias Unberath arxiv

Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets with precise spatial locations and temp…

Reinforcement Learning

Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing

2025-10-31 · Yijia Wang, Yiqing Shen, Weiming Chen, Zhihai He arxiv

Existing image editing methods can handle simple editing instructions very well. To deal with complex editing instructions, they often need to jointly fine-tune the large language models (LLMs) and diffusion models (DMs)…

Image Editing