paper-with-me

홈 › Papers

Policy Optimization Prefers The Path of Least Resistance

2025-10-22 · Debdeep Sanyal, Aakash Sen Sharma, Dhruv Kumar, Saurabh Deshpande, Murari Mandal arxiv

Policy optimization (PO) algorithms are used to refine Large Language Models for complex, multi-step reasoning. Current state-of-the-art pipelines enforce a strict think-then-answer format to elicit chain-of-thought (CoT); however, the behavior of PO when these rigid constraints are relaxed into an open-ended CoT structure remains an under-studied question. We investigate this gap with an extensive suite of controlled experiments and identify a consistent principle: \textit{policy optimization consistently follows the path of least resistance}. When afforded the flexibility to interleave reasoning and response, policy optimization consistently learns to discard explicit reasoning, causing the policy to degenerate to a direct \texttt{<answer>}-only format. This outcome holds true across various models and algorithms. We find that this collapse in format is persistent even when the complex \texttt{<think><answer>} format is assigned up to 4x larger reward weights. We formalize this principle through a series of controlled reward decomposition experiments, demonstrating a clear hierarchy: PO systematically optimizes for the simplest reward component first, a preference that holds even when faced with mutually exclusive choices or strong incentives for more complex behaviors. Finally, we show that successful convergence on the high-reward shortcut is not a low-effort drift but is driven by the optimization process that requires the KL-regularized policy to have sufficient freedom to make a significant shift from its initial prior. Our findings reveal that granting policies the freedom to diverge is a double-edged sword: while necessary for discovering high-reward shortcuts, it also creates a powerful incentive to game the simplest aspects of the reward function, posing a critical challenge for reward hacking under alignment.

📄 PDF Abstract BibTeX arXiv:2510.21853

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Implementability of Liberalism

2025-06-19 · Héctor Hermida-Rivera

This note shows that under the unrestricted domain, there exists a choice liberal and Nash implementable social choice rule if and only if there are at least three players and the outcome set is at least twice as large a…

The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus

2026-01-29 · Ishan Jindal, Sai Prashanth Akuthota, Jayant Taneja, Sachin Dev Sharma arxiv

Large language models achieve strong reasoning performance, but inference strategies such as Self-Consistency (SC) are computationally expensive, as they fully expand all reasoning traces. We introduce PoLR (Path of Leas…

Decoding Microbial Enigmas: Unleashing the Power of Artificial Intelligence in Analyzing Antibiotic-Resistant Pathogens and their Impact on Human Health

2023-07-27 · Maitham G. Yousif

In this research, medical information from 1200 patients across various hospitals in Iraq was collected over a period of 3 years, from February 3, 2018, to March 5, 2021. The study encompassed several infections, includi…

Design-Technology Co-Optimization for NVM-based Neuromorphic Processing Elements

2022-03-10 · Shihao Song, Adarsha Balaji, Anup Das, Nagarajan Kandasamy

Neuromorphic hardware platforms can significantly lower the energy overhead of a machine learning inference task. We present a design-technology tradeoff analysis to implement such inference tasks on the processing eleme…

BIG-bench Machine Learning

Crowd Prefers the Middle Path: A New IAA Metric for Crowdsourcing Reveals Turker Biases in Query Segmentation

2013-08-01 · ACL 2013 8 · Rohan Ramanath, Monojit Choudhury, Kalika Bali, Rishiraj Saha Roy
Chunking