paper-with-me

Papers

Language Models use Lookbacks to Track Beliefs

2025-05-20 · Nikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl, Yonatan Belinkov, Tamar Rott Shaham, David Bau, Atticus Geiger

How do language models (LMs) represent characters' beliefs, especially when those beliefs may differ from reality? This question lies at the heart of understanding the Theory of Mind (ToM) capabilities of LMs. We analyze Llama-3-70B-Instruct's ability to reason about characters' beliefs using causal mediation and abstraction. We construct a dataset that consists of simple stories where two characters each separately change the state of two objects, potentially unaware of each other's actions. Our investigation uncovered a pervasive algorithmic pattern that we call a lookback mechanism, which enables the LM to recall important information when it becomes necessary. The LM binds each character-object-state triple together by co-locating reference information about them, represented as their Ordering IDs (OIs) in low rank subspaces of the state token's residual stream. When asked about a character's beliefs regarding the state of an object, the binding lookback retrieves the corresponding state OI and then an answer lookback retrieves the state token. When we introduce text specifying that one character is (not) visible to the other, we find that the LM first generates a visibility ID encoding the relation between the observing and the observed character OIs. In a visibility lookback, this ID is used to retrieve information about the observed character and update the observing character's beliefs. Our work provides insights into the LM's belief tracking mechanisms, taking a step toward reverse-engineering ToM reasoning in LMs.

📄 PDF Abstract BibTeX arXiv:2505.14685

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dynamic Theory of Mind as a Temporal Memory Problem: Evidence from Large Language Models

2026-03-15 · Thuy Ngoc Nguyen, Duy Nhat Phan, Cleotilde Gonzalez arxiv

Theory of Mind (ToM) is central to social cognition and human-AI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment…

Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs

2022-11-15 · Kyle Richardson, Ronen Tamari, Oren Sultan, Reut Tsarfaty 외

Can we teach natural language understanding models to track their beliefs through intermediate points in text? We propose a representation learning framework called breakpoint modeling that allows for learning of this ty…

Natural Language UnderstandingRelational ReasoningRepresentation Learning

Classification-Aided Robust Multiple Target Tracking Using Neural Enhanced Message Passing

2023-10-19 · Xianglong Bai, Zengfu Wang, Quan Pan, Tao Yun 외

We address the challenge of tracking an unknown number of targets in strong clutter environments using measurements from a radar sensor. Leveraging the range-Doppler spectra information, we identify the measurement class…

A General Approach for Lookback Option Pricing under Markov Models

2021-12-01 · Gongqiu Zhang, Lingfei Li

We propose a very efficient method for pricing various types of lookback options under Markov models. We utilize the model-free representations of lookback option prices as integrals of first passage probabilities. We co…

Evaluating Theory of Mind in Question Answering

2018-08-28 · EMNLP 2018 10 · Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik 외

We propose a new dataset for evaluating question answering models with respect to their capacity to reason about beliefs. Our tasks are inspired by theory-of-mind experiments that examine whether children are able to rea…

Question Answering