paper-with-me

홈 › Papers

Chain and Causal Attention for Efficient Entity Tracking

2024-10-07 · Erwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre Allauzen

This paper investigates the limitations of transformers for entity-tracking tasks in large language models. We identify a theoretical constraint, showing that transformers require at least $\log_2 (n+1)$ layers to handle entity tracking with $n$ state changes. To address this issue, we propose an efficient and frugal enhancement to the standard attention mechanism, enabling it to manage long-term dependencies more efficiently. By considering attention as an adjacency matrix, our model can track entity states with a single layer. Empirical results demonstrate significant improvements in entity tracking datasets while keeping competitive performance on standard natural language modeling. Our modified attention allows us to achieve the same performance with drastically fewer layers. Additionally, our enhanced mechanism reveals structured internal representations of attention. Extensive experiments on both toy and complex datasets validate our approach. Our contributions include theoretical insights, an improved attention mechanism, and empirical validation.

📄 PDF Abstract BibTeX arXiv:2410.05565

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Causal Reasoning of Entities and Events in Procedural Texts

2023-01-26 · Li Zhang, Hainiu Xu, Yue Yang, Shuyan Zhou 외

Entities and events are crucial to natural language reasoning and common in procedural texts. Existing work has focused either exclusively on entity state tracking (e.g., whether a pan is hot) or on event reasoning (e.g.…

Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

2020-07-29 · ECCV 2020 8 · Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu 외

Existing Multiple-Object Tracking (MOT) methods either follow the tracking-by-detection paradigm to conduct object detection, feature extraction and data association separately, or have two of the three subtasks integrat…

Multiple Object TrackingObjectobject-detectionObject Detection+2

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought

2026-06-14 · Yaoting Huang, Yifu Yuan, Linqi Han, Chengwen Li 외 arxiv

Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models r…

Spatial ReasoningVisual Grounding

CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks

2025-07-18 · Yanan Wang, Julio Vizcarra, Zhi Li, Hao Niu 외 arxiv

Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fine-grained object-level video understand…

Temporal Relation Extraction

A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models

2026-06-07 · Soyoung Oh, Vera Demberg arxiv

To interpret context correctly and retrieve relevant information, large language models must bind entities to their attributes and update these bindings as state changes. We analyze how LLMs implement this binding proces…