paper-with-me

홈 › Papers

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

2026-02-09 · Raghu Arghal, Fade Chen, Niall Dalton, Evgenii Kortukov, Calum McNamara, Angelos Nalmpantis, Moksh Nirvaan, Gabriele Sarti, Mario Giulianelli arxiv

Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models' internal representations. As a case study, we examine an LLM agent navigating a 2D grid world towards a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues towards immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.

📄 PDF Abstract BibTeX arXiv:2602.08964

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating the Goal-Directedness of Large Language Models

2025-04-16 · Tom Everitt, Cristina Garbacea, Alexis Bellot, Jonathan Richens 외

To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require information gathering, cognitive effort, a…

Towards Measuring Goal-Directedness in AI Systems

2024-10-07 · Dylan Xu, Juan-Pablo Rivera

Recent advances in deep learning have brought attention to the possibility of creating advanced, general AI systems that outperform humans across many tasks. However, if these systems pursue unintended goals, there could…

Reinforcement Learning (RL)

Measuring Goal-Directedness

2024-12-06 · Matt MacDermott, James Fox, Francesco Belardinelli, Tom Everitt

We define maximum entropy goal-directedness (MEG), a formal measure of goal-directedness in causal models and Markov decision processes, and give algorithms for computing it. Measuring goal-directedness is important, as …

Goal-Directedness is in the Eye of the Beholder

2025-08-18 · Nina Rajcic, Anders Søgaard arxiv

Our ability to predict the behavior of complex agents turns on the attribution of goals. Probing for goal-directed behavior comes in two flavors: Behavioral and mechanistic. The former proposes that goal-directedness can…

Realistic honeypot evaluations for scheming propensity

2026-05-28 · Victoria Krakovna, David Lindner, Lewis Ho, Sebastian Farquhar 외 arxiv

We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming honeypot evaluations take the form of coding tasks in Google's alig…