paper-with-me

홈 › Papers

Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors

2026-03-14 · Mark Rofin, Jalal Naghiyev, Michael Hahn arxiv

Trained Transformers have been shown to compute abstract features that appear redundant for predicting the immediate next token. We identify which components of the gradient signal from the next-token prediction objective give rise to this phenomenon, and we propose a method to estimate the influence of those components on the emergence of specific features. After validating our approach on toy tasks, we use it to interpret the origins of the world model in OthelloGPT and syntactic features in a small language model. Finally, we apply our framework to a pretrained LLM, showing that features with extremely high or low influence on future tokens tend to be related to formal reasoning domains such as code. Overall, our work takes a step toward understanding hidden features of Transformers through the lens of their development during training.

📄 PDF Abstract BibTeX arXiv:2603.14087

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Proof of a perfect platonic representation hypothesis

2025-07-01 · Liu Ziyin, Isaac Chuang arxiv

In this note, we elaborate on and explain in detail the proof given by Ziyin et al. (2025) of the ``perfect" Platonic Representation Hypothesis (PRH) for the embedded deep linear network model (EDLN). We show that if tra…

Representation Learning

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

2026-02-05 · John Yan, Michael Yu, Yuqi Sun, Alexander Duffy 외 arxiv

Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have rec…

Multi-agent Reinforcement Learning

The Safety Filter: A Unified View of Safety-Critical Control in Autonomous Systems

2023-09-11 · Kai-Chieh Hsu, Haimin Hu, Jaime Fernández Fisac

Recent years have seen significant progress in the realm of robot autonomy, accompanied by the expanding reach of robotic technologies. However, the emergence of new deployment domains brings unprecedented challenges in …

NExT-Chat: An LMM for Chat, Detection and Segmentation

2023-11-08 · Ao Zhang, Yuan YAO, Wei Ji, Zhiyuan Liu 외

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance the level of visual comprehensio…

Referring ExpressionReferring Expression SegmentationVisual Grounding

Emergence in artificial life

2021-04-30 · Carlos Gershenson

Even when concepts similar to emergence have been used since antiquity, we lack an agreed definition. However, emergence has been identified as one of the main features of complex systems. Most would agree on the stateme…

Artificial Life