paper-with-me

홈 › Papers

Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium

2025-08-28 · Christopher R. Lee-Jenkins arxiv

Decoding in large language models is often described as scoring tokens and normalizing with softmax. We give a minimal, self-contained account of this step as a constrained variational principle on the probability simplex. The discrete, normalization-respecting ascent is the classical multiplicative-weights (entropic mirror) update; its continuous-time limit is the replicator flow. From these ingredients we prove that, for a fixed context and temperature, the next-token distribution follows a smooth trajectory inside the simplex and converges to the softmax equilibrium. This formalizes the common ``manifold traversal'' intuition at the output-distribution level. The analysis yields precise, practice-facing consequences: temperature acts as an exact rescaling of time along the same trajectory, while top-k and nucleus sampling restrict the flow to a face with identical guarantees. We also outline a controlled account of path-dependent score adjustments and their connection to loop-like, hallucination-style behavior. We make no claims about training dynamics or internal representations; those are deferred to future work.

📄 PDF Abstract BibTeX arXiv:2508.21186

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Humanoid Locomotion as Next Token Prediction

2024-02-29 · Ilija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran 외

We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive prediction of sensorimotor trajectories. …

Humanoid ControlPrediction

Lines of Thought in Large Language Models

2024-10-02 · Raphaël Sarfati, Toni J. B. Liu, Nicolas Boullé, Christopher J. Earls

Large Language Models achieve next-token prediction by transporting a vectorized piece of text (prompt) across an accompanying embedding space under the action of successive transformer layers. The resulting high-dimensi…

In-Context Imitation Learning via Next-Token Prediction

2024-08-28 · Letian Fu, Huang Huang, Gaurav Datta, Lawrence Yunliang Chen 외

We explore how to enhance next-token prediction models to perform in-context imitation learning on a real robot, where the robot executes new tasks by interpreting contextual information provided during the input phase, …

Imitation LearningPrediction

Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation

2025-10-10 · Yao Teng, Fuyun Wang, Xian Liu, Zhekai Chen 외 arxiv

As a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to…

Text-to-Image Generation

Scaling Next-Brain-Token Prediction for MEG

2026-01-28 · Richard Csaky arxiv

We present a large autoregressive model for source-space MEG that scales next-token prediction to long context across datasets and scanners: handling a corpus of over 500 hours and thousands of sessions across the three …