paper-with-me

Papers

Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization

2026-06-29 · Kunal Samanta, Ari Holtzman, Peter West arxiv

The alignment of language models is typically studied through the lens of capability benchmarks, but the dynamics of how models change during post-training remain poorly understood. We argue that the physical sciences, and thermodynamic phase-transition theory in particular, offer a principled and underexplored vocabulary for reasoning about these dynamics. As a case study, we instantiate this position through the lens of material Crystallization, which is a well-studied thermodynamic phase transition. For tasks like random number generation, this breaks into 3 phases: (1) the high entropy liquid phase in the pretrained model, with many distinct sampling distributions promptable from the model; (2) the nucleation phase caused by supervised finetuning, in which behavior collapses onto a single seed distribution present in the pretrained LLM; and (3) a settling phase in which reinforcement learning techniques redistribute probability of the collapsed distribution, but largely keep it concentrated on the same options as the seed distribution. We propose intuitive metrics to verify the transitions between these phases, and validate the idea across a range of random tasks. Crystallization is one instance of a broader class of physical frameworks we believe alignment research should import to answer questions about where alignment-induced structure comes from, why it converges where it does, and what it fundamentally cannot change.

📄 PDF Abstract BibTeX arXiv:2606.29933

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Can Vision Language Models Learn Intuitive Physics from Interaction?

2026-02-05 · Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric Schulz arxiv

Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned model…

Reinforcement Learning

MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment Tasks

2023-10-30 · NeurIPS 2023 11 · Allen Nie, Yuhui Zhang, Atharva Amdekar, Chris Piech 외

Human commonsense understanding of the physical and social world is organized around intuitive theories. These theories support making causal and moral judgments. When something bad happens, we naturally ask: who did wha…

Language ModelingLanguage Modelling

Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires

2025-07-14 · Simon Münker arxiv

Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their lin…

Modeling human intuitions about liquid flow with particle-based simulation

2018-09-05 · Christopher J. Bates, Ilker Yildirim, Joshua B. Tenenbaum, Peter Battaglia

Humans can easily describe, imagine, and, crucially, predict a wide variety of behaviors of liquids--splashing, squirting, gushing, sloshing, soaking, dripping, draining, trickling, pooling, and pouring--despite tremendo…

Scene Understanding

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

2025-12-05 · Zijun Wang, Panwen Hu, Jing Wang, Terry Jingchen Zhang 외 arxiv

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-sca…

Video Generation