The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
Direct Preference Optimization (DPO) is often tuned as if increasing alignment pressure (controlled by $β$) yields progressively "better" behavior. We instead treat $β$ as a control parameter and densely sweep it for three 7B open-weight families under a fixed DPO recipe. In Mistral, capability is sharply non-monotonic: aggregated logic-probe margins become positive only in a narrow band near $β\approx 10^{-2}$ and revert outside it, with boundary points that are seed-sensitive. Across architectures under the same sweep, we observe qualitatively different response modes: sharp reorganization in Mistral, selective changes in Llama, and smooth trade-offs in Qwen. Critically, the DPO preference margin can anticorrelate with reasoning capability (Pearson $r=-0.91$ for Llama logic), so margin-based selection can prefer capability-impaired models. Training path also matters: exposure to high $β$ induces capability losses that persist even after $β$ is reduced (hysteresis). These findings motivate capability-resolved evaluation across the $β$ landscape rather than reliance on margins or aggregate benchmarks.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Latent-Hysteresis Graph ODEs: Modeling Coupled Topology-Feature Evolution via Continuous Phase Transitions
Graph neural ordinary differential equations (Graph ODEs) extend graph learning from discrete message-passing layers to continuous-time representation flows. While it supports adaptive long-range propagation, we show tha…
Graph LearningOutward Sodium Current, Fine-structure Constant and Ferroelectric Hysteresis Regimes in the Giant Squid Axon Propagating Action Potential: a Phase Space Approach
We derive a charge-conserving phase space cable equation for the propagating action potential, expressing ionic, membrane, and capacitive currents as functions of membrane potential. Analysis of the ionic current during …
A phase-field model for active contractile surfaces
The morphogenesis of cells and tissues involves an interplay between chemical signals and active forces on their surrounding surface layers. The complex interaction of hydrodynamics and material flows on such active surf…
Contact mechanicsmodelNoise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks
Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature. Below a critical regularization str…
Some elementary mechanisms for critical transitions and hysteresis in simple predator prey models
Trait-mediated indirect effects are increasingly acknowledged as important components in the dynamics of ecological systems. The hamiltonian form of the LV equations is traditionally modified by adding density dependence…