paper-with-me

홈 › Papers

Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs

2025-11-03 · Bo Bai arxiv

Despite the empirical successes of Large Language Models (LLMs), the prevailing paradigm is heuristic and experiment-driven, tethered to massive compute and data, while a first-principles theory remains absent. This treatise develops a Semantic Information Theory at the confluence of statistical physics, signal processing, and classical information theory, organized around a single paradigm shift: replacing the classical BIT - a microscopic substrate devoid of semantic content - with the macroscopic TOKEN as the atomic carrier of meaning and reasoning. Within this framework we recast attention and the Transformer as energy-based models, and interpret semantic embedding as vectorization on the semantic manifold. Modeling the LLM as a stateful channel with feedback, we adopt Massey's directed information as the native causal measure of autoregressive generation, from which we derive a *directed rate-distortion function for pre-training, a directed rate-reward function for RL-based post-training, and a sub-martingale account of inference-time semantic information flow. This machinery makes precise the identification of next-token prediction with Granger causal inference, and sharpens the limits of LLM reasoning against Pearl's Ladder of Causation - affirming that *whereas the BIT defined the Information Epoch, the TOKEN will define the AI Epoch.

📄 PDF Abstract BibTeX arXiv:2511.01202

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting

2024-05-28 · Suraj Anand, Michael A. Lepori, Jack Merullo, Ellie Pavlick

Language models have the ability to perform in-context learning (ICL), allowing them to flexibly adapt their behavior based on context. This contrasts with in-weights learning, where information is statically encoded in …

In-Context Learning

Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance

2026-06-04 · Gizem Yüce, Giorgos Nikolaou, Nicolas Flammarion arxiv

Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. For autoregressive language models, not all tokens in a forget sample are equally relevant to forgettin…

What Are You Token About? Dense Retrieval as Distributions Over the Vocabulary

2022-12-20 · Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov 외

Dual encoders are now the dominant architecture for dense retrieval. Yet, we have little understanding of how they represent text, and why this leads to good performance. In this work, we shed light on this question via …

Retrieval

Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning

2025-08-06 · Ali Taheri, Alireza Taban, Qizhou Wang, Shanshan Ye 외 arxiv

Supervised fine-tuning (SFT) plays a critical role for pretrained large language models (LLMs), notably enhancing their capacity to acquire domain-specific knowledge while preserving or potentially augmenting their gener…

The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives

2019-09-03 · IJCNLP 2019 11 · Elena Voita, Rico Sennrich, Ivan Titov

We seek to understand how the representations of individual tokens and the structure of the learned feature space evolve between layers in deep neural networks under different learning objectives. We focus on the Transfo…

Language ModelingLanguage ModellingMachine TranslationMasked Language Modeling+1