paper-with-me

홈 › Papers

Mutual Information Scaling and Expressive Power of Sequence Models

2019-05-10 · Huitao Shen

Sequence models assign probabilities to variable-length sequences such as natural language texts. The ability of sequence models to capture temporal dependence can be characterized by the temporal scaling of correlation and mutual information. In this paper, we study the mutual information of recurrent neural networks (RNNs) including long short-term memories and self-attention networks such as Transformers. Through a combination of theoretical study of linear RNNs and empirical study of nonlinear RNNs, we find their mutual information decays exponentially in temporal distance. On the other hand, Transformers can capture long-range mutual information more efficiently, making them preferable in modeling sequences with slow power-law mutual information, such as natural languages and stock prices. We discuss the connection of these results with statistical mechanics. We also point out the non-uniformity problem in many natural language datasets. We hope this work provides a new perspective in understanding the expressive power of sequence models and shed new light on improving the architecture of them.

📄 PDF Abstract BibTeX arXiv:1905.04271

Code (1)

openai/gpt-2 공식 구현 tf

Similar Papers 제목 키워드 기반

On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding

2024-10-02 · Kevin Xu, Issei Sato

Looped Transformers provide advantages in parameter efficiency, computational capabilities, and generalization for reasoning tasks. However, their expressive power regarding function approximation remains underexplored. …

GRAB: An LLM-Inspired Sequence-First Click-Through Rate Prediction Modeling Paradigm

2026-02-02 · Shaopeng Chen, Chuyue Xie, Huimin Ren, Shaozong Zhang 외 arxiv

Traditional Deep Learning Recommendation Models (DLRMs) face increasing bottlenecks in performance and efficiency, often struggling with generalization and long-sequence modeling. Inspired by the scaling success of Large…

Click-Through Rate Prediction

Transformers are Efficient Compilers, Provably

2024-10-07 · Xiyu Zhai, Runlong Zhou, Liao Zhang, Simon Shaolei Du

Transformer-based large language models (LLMs) have demonstrated surprisingly robust performance across a wide range of language-related tasks, including programming language understanding and generation. In this paper, …

L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling

2025-03-06 · Zhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo 외

We rigorously establish a bipartite mutual information scaling law in natural language that governs long-range dependencies. This scaling law, which we show is distinct from and scales independently of the conventional t…

Language ModelingLanguage ModellingState Space Models

Tensor networks and efficient descriptions of classical data

2021-03-11 · Sirui Lu, Márton Kanász-Nagy, Ivan Kukuljan, J. Ignacio Cirac

We investigate the potential of tensor network based machine learning methods to scale to large image and text data sets. For that, we study how the mutual information between a subregion and its complement scales with t…

Tensor Networks