paper-with-me

홈 › Papers

Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap

2025-10-31 · Chun-Hao Yang, Bo-Han Feng, Tzu-Yuan Lai, Yan Yu Chen, Yin-Kai Dean Huang, Shou-De Lin arxiv

Optimizing training performance in large language models (LLMs) remains an essential challenge, particularly in improving model performance while maintaining computational costs. This work challenges the conventional approach of training LLMs using next-token prediction (NTP), arguing that by predicting information-rich tokens during training, there is a more effective way to train LLMs. We investigate the impact of the proposed solution in three kinds of tasks for LLMs: arithmetic, multi-label classification of text, and natural-language generation. This work offers a principled approach to optimizing LLM training, advancing both model performance and theoretical understanding of the target-token selection strategies.

📄 PDF Abstract BibTeX arXiv:2511.00198

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Label Classification

Similar Papers 제목 키워드 기반

Beyond Next Token Prediction: Patch-Level Training for Large Language Models

2024-07-17 · Chenze Shao, Fandong Meng, Jie zhou

The prohibitive training costs of Large Language Models (LLMs) have emerged as a significant bottleneck in the development of next-generation LLMs. In this paper, we show that it is possible to significantly reduce the t…

Language ModelingLanguage Modelling

Emu3: Next-Token Prediction is All You Need

2024-09-27 · Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun 외

While next-token prediction is considered a promising path towards artificial general intelligence, it has struggled to excel in multimodal tasks, which are still dominated by diffusion models (e.g., Stable Diffusion) an…

AllImage GenerationPrediction+2

Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries

2025-10-16 · Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki 외 arxiv

Next-token prediction (NTP) has driven the success of large language models (LLMs), but it struggles with long-horizon reasoning, planning, and creative writing, with these limitations largely attributed to teacher-force…

LLMs are Not Just Next Token Predictors

2024-08-06 · Stephen M. Downes, Patrick Forber, Alex Grzankowski

LLMs are statistical models of language learning through stochastic gradient descent with a next token prediction objective. Prompting a popular view among AI modelers: LLMs are just next token predictors. While LLMs are…

Prediction

Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators

2025-04-15 · Phill Kyu Rhee

Large Language Models (LLMs), powered by Transformers, have demonstrated human-like intelligence capabilities, yet their underlying mechanisms remain poorly understood. This paper presents a novel framework for interpret…