paper-with-me

홈 › Papers

Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries

2025-10-16 · Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki, Ioannis Mitliagkas, David Lopez-Paz, Kartik Ahuja arxiv

Next-token prediction (NTP) has driven the success of large language models (LLMs), but it struggles with long-horizon reasoning, planning, and creative writing, with these limitations largely attributed to teacher-forced training. Multi-token prediction (MTP) partially mitigates these issues by predicting several future tokens at once, but it mostly captures short-range dependencies and offers limited improvement. We propose future summary prediction (FSP), which trains an auxiliary head to predict a compact representation of the long-term future, preserving information relevant for long-form generations. We explore two variants of FSP: handcrafted summaries, for example, a bag of words summary of the future sequence, and learned summaries, which use embeddings produced by a reverse language model trained from right-to-left order. Large-scale pretraining experiments (3B and 8B-parameter models) demonstrate that FSP provides improvements over both NTP and MTP across math, reasoning, and coding benchmarks.

📄 PDF Abstract BibTeX arXiv:2510.14751

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs

2024-09-23 · Clément Christophe, Tathagata Raha, Svetlana Maslenkova, Muhammad Umar Salman 외

Large Language Models (LLMs) have demonstrated significant potential in transforming clinical applications. In this study, we investigate the efficacy of four techniques in adapting LLMs for clinical use-cases: continuou…

Prompt Engineering

PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling

2025-10-11 · Guilin Li, Yun Zhang, Xiuyuan Chen, Chengqi Li 외 arxiv

Large language models (LLMs) have shown that generative pretraining can distill vast world knowledge into compact token representations. While LLMs encapsulate extensive world knowledge, they remain limited in modeling t…

Representation LearningFraud Detection

Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data

2026-01-16 · Minseo Kwak, Jaehyung Kim arxiv

The opacity of massive pretraining corpora in Large Language Models (LLMs) raises significant privacy and copyright concerns, making pretraining data detection a critical challenge. Existing state-of-the-art methods typi…

One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

2026-01-31 · Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A. Lee 외 arxiv

Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages. The sequenti…

Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension

2024-04-13 · MengNan Qi, Yufan Huang, Yongqiang Yao, Maoquan Wang 외

Large language models (LLMs) has experienced exponential growth, they demonstrate remarkable performance across various tasks. Notwithstanding, contemporary research primarily centers on enhancing the size and quality of…

Code CompletionSentenceSentence EmbeddingSentence-Embedding