paper-with-me

Papers

Predicting Future Behaviors in Reasoning Models Enables Better Steering

2026-06-09 · Evgenii Kortukov, Piotr Komorowski, Florian Klein, Paula Engl, Gabriele Sarti, Seong Joon Oh, Sebastian Lapuschkin, Wojciech Samek arxiv

Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output quality. We argue that prior steering work implicitly relies on internal features that detect behavior in already generated text. We show that these detection features are poor predictors of future behavioral outcomes, and thus not the natural intervention target. Instead, we train activation probes to predict future behavior likelihoods from intermediate reasoning steps. These probes predict the most likely behavior with 64%-91% accuracy, revealing a separate type of internal prediction features. Building on these prediction features, we introduce a text-level steering method, Future Probe Controlled Generation. FPCG samples multiple candidate sentences and chooses the best one according to a probe predicting the future behavior likelihood. This enables steering with almost no output quality degradation. FPCG also enables steering in several evaluations where activation steering fails. These results show that distinguishing detection and prediction features enables a more nuanced approach to controlling LRM behaviors.

📄 PDF Abstract BibTeX arXiv:2606.11172

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TrendGNN: Towards Understanding of Epidemics, Beliefs, and Behaviors

2025-11-29 · Mulin Tian, Ajitesh Srivastava arxiv

Epidemic outcomes have a complex interplay with human behavior and beliefs. Most of the forecasting literature has focused on the task of predicting epidemic signals using simple mechanistic models or black-box models, s…

Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors

2025-09-16 · Aniket Didolkar, Nicolas Ballas, Sanjeev Arora, Anirudh Goyal arxiv

Large language models (LLMs) now solve multi-step problems by emitting extended chains of thought. During the process, they often re-derive the same intermediate steps across problems, inflating token usage and latency. …

RAG-based Explainable Prediction of Road Users Behaviors for Automated Driving using Knowledge Graphs and Large Language Models

2024-05-01 · Mohamed Manzour Hussien, Angie Nataly Melo, Augusto Luis Ballardini, Carlota Salinas Maldonado 외

Prediction of road users' behaviors in the context of autonomous driving has gained considerable attention by the scientific community in the last years. Most works focus on predicting behaviors based on kinematic inform…

Autonomous DrivingBayesian InferenceKnowledge Graph EmbeddingsKnowledge Graphs+3

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

2025-05-23 · Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu 외

Visual language models (VLMs) have attracted increasing interest in autonomous driving due to their powerful reasoning capabilities. However, existing VLMs typically utilize discrete text Chain-of-Thought (CoT) tailored …

Autonomous DrivingImage GenerationTrajectory PlanningVisual Reasoning

Scenario-Transferable Semantic Graph Reasoning for Interaction-Aware Probabilistic Prediction

2020-04-07 · Yeping Hu, Wei Zhan, Masayoshi Tomizuka

Accurately predicting the possible behaviors of traffic participants is an essential capability for autonomous vehicles. Since autonomous vehicles need to navigate in dynamically changing environments, they are expected …

Autonomous DrivingAutonomous VehiclesNavigate