paper-with-me

Papers

Large language models transition from integrating across position-yoked, exponential windows to structure-yoked, power-law windows

2023-09-21 · NeurIPS 2023 11

Modern language models excel at integrating across long temporal scales needed to encode linguistic meaning and show non-trivial similarities to biological neural systems. Prior work suggests that human brain responses to language exhibit hierarchically organized "integration windows" that substantially constrain the overall influence of an input token (e.g., a word) on the neural response. However, little prior work has attempted to use integration windows to characterize computations in large language models (LLMs). We developed a simple word-swap procedure for estimating integration windows from black-box language models that does not depend on access to gradients or knowledge of the model architecture (e.g., attention weights). Using this method, we show that trained LLMs exhibit stereotyped integration windows that are well-fit by a convex combination of an exponential and a power-law function, with a partial transition from exponential to power-law dynamics across network layers. We then introduce a metric for quantifying the extent to which these integration windows vary with structural boundaries (e.g., sentence boundaries), and using this metric, we show that integration windows become increasingly yoked to structure at later network layers. None of these findings were observed in an untrained model, which as expected integrated uniformly across its input. These results suggest that LLMs learn to integrate information in natural language using a stereotyped pattern: integrating across position-yoked, exponential windows at early layers, followed by structure-yoked, power-law windows at later layers. The methods we describe in this paper provide a general-purpose toolkit for understanding temporal integration in language models, facilitating cross-disciplinary research at the intersection of biological and artificial intelligence.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking Compositionality with Formal Languages

2022-08-17 · COLING 2022 10 · Josef Valvoda, Naomi Saphra, Jonathan Rawski, Adina Williams 외

Recombining known primitive concepts into larger novel combinations is a quintessentially human cognitive capability. Whether large neural models in NLP can acquire this ability while learning from data is an open questi…

BenchmarkingOpen-Ended Question Answering

H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model

2026-02-11 · Jinbang Huang, Wenyuan Chen, Zhiyuan Li, Oscar Pang 외 arxiv

World models are becoming central to robotic planning and control as they enable prediction of future state transitions. Existing approaches often emphasize video generation or natural-language prediction, which are diff…

Visual GroundingVideo GenerationMotion Planning

VRoPE: Rotary Position Embedding for Video Large Language Models

2025-02-17 · Zikang Liu, Longteng Guo, Yepeng Tang, Tongtian Yue 외

Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Exist…

PositionVideo Understanding

Highway Graph to Accelerate Reinforcement Learning

2024-05-20 · Zidu Yin, Zhen Zhang, Dong Gong, Stefano V. Albrecht 외

Reinforcement Learning (RL) algorithms often struggle with low training efficiency. A common approach to address this challenge is integrating model-based planning algorithms, such as Monte Carlo Tree Search (MCTS) or Va…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

A Closer Look at Reward Decomposition for High-Level Robotic Explanations

2023-04-25 · Wenhao Lu, Xufeng Zhao, Sven Magg, Martin Gromniak 외

Explaining the behaviour of intelligent agents learned by reinforcement learning (RL) to humans is challenging yet crucial due to their incomprehensible proprioceptive states, variational intermediate goals, and resultan…

Reinforcement Learning (RL)Vocal Bursts Intensity Prediction