paper-with-me

홈 › Papers

A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners

2026-06-02 · Patrick Emami, Nan Qiang, Peter Graf arxiv

Supervised fine-tuning (SFT) improves end-to-end classical planning in large language models (LLMs), but do these models also learn to represent and reason about the planning problems they are solving? Due to the relative complexity of classical planning problems and the challenge that end-to-end plan generation poses for LLMs, it has been difficult to explore this question. In our work, we devise and perform a series of interpretability experiments that holistically interrogate world model recovery by examining both internal representations and generative capabilities of fine-tuned LLMs. We find that: a) Supervised fine-tuning on valid action sequences enables LLMs to linearly encode action validity and some state predicates. b) Models that struggle to use output probabilities for classifying action validity may still learn internal representations that separate valid from invalid actions. c) Broader state space coverage during fine-tuning, such as from random walk data, yields more accurate recovery of the underlying world model. In summary, this work contributes a recipe for applying interpretability techniques to planning LLMs and generates insights that shed light on open questions about how knowledge is represented in LLMs.

📄 PDF Abstract BibTeX arXiv:2606.03685

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

2026-05-18 · Wenhao Shen, Ming Zhou, Hengyuan Zhang, Siyuan Bian 외 arxiv

Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expressive whole-body motion are required. Existing video HMR methods prod…

Human Mesh Recovery

Single View Seafloor Recovery from Imaging Sonar via Differentiable Rendering

2026-05-22 · Sevan Brodjian, Michael Hobley, Pietro Perona arxiv

Sonar is often the only modality suitable for high-resolution imaging underwater due to light attenuation and turbidity. Forward-looking imaging sonar provides measurements over range and horizontal angle but collapses v…

From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning

2026-08-01 · Hao Yuan, Yuxin Wang, Lei Ji, Zhiwei Yu arxiv

Physical-world interaction is inherently dynamic, as environments can evolve during execution, requiring agents to adapt their plans under non-stationary conditions. We study this challenge through long-horizon embodied …

EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence

2026-06-10 · Xin Zhou, Cong Miao arxiv

In this paper, we propose the Enhanced World Action Model (EWAM), a closed-loop online adaptation architecture built upon a pretrained and fully frozen Cosmos3 backbone network. Evaluated entirely under a zero-shot task …

Anomaly Detection

Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis

2026-03-13 · Dayou Li, Jiuzhou Lei, Hao Wang, Lulin Liu 외 arxiv

While recent foundation models have significantly advanced robotic manipulation, these systems still struggle to autonomously recover from execution errors. Current failure-learning paradigms rely on either costly and un…