paper-with-me

홈 › Papers

Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension

2024-04-13 · MengNan Qi, Yufan Huang, Yongqiang Yao, Maoquan Wang, Bin Gu, Neel Sundaresan

Large language models (LLMs) has experienced exponential growth, they demonstrate remarkable performance across various tasks. Notwithstanding, contemporary research primarily centers on enhancing the size and quality of pretraining data, still utilizing the next token prediction task on autoregressive transformer model structure. The efficacy of this task in truly facilitating the model's comprehension of code logic remains questionable, we speculate that it still interprets code as mere text, while human emphasizes the underlying logical knowledge. In order to prove it, we introduce a new task, "Logically Equivalent Code Selection," which necessitates the selection of logically equivalent code from a candidate set, given a query code. Our experimental findings indicate that current LLMs underperform in this task, since they understand code by unordered bag of keywords. To ameliorate their performance, we propose an advanced pretraining task, "Next Token Prediction+". This task aims to modify the sentence embedding distribution of the LLM without sacrificing its generative capabilities. Our experimental results reveal that following this pretraining, both Code Llama and StarCoder, the prevalent code domain pretraining models, display significant improvements on our logically equivalent code selection task and the code completion task.

📄 PDF Abstract BibTeX arXiv:2404.08885

Code (0)

등록된 구현이 없습니다.

Tasks

Code CompletionSentenceSentence EmbeddingSentence-Embedding

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

The pitfalls of next-token prediction

2024-03-11 · Gregor Bachmann, Vaishnavh Nagarajan

Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective. As a starting poin…

MambaMisconceptionsPrediction

Next-token prediction capacity: general upper bounds and a lower bound for transformers

2024-05-22 · Liam Madden, Curtis Fox, Christos Thrampoulidis

Given a sequence of tokens, such as words, the task of next-token prediction is to predict the next-token conditional probability distribution. Decoder-only transformers have become effective models for this task, but th…

DecoderMemorization

NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training

2024-09-13 · Minglun Han, Ye Bai, Chen Shen, Youjia Huang 외

Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT and BEST-RQ, focus on utilizing non-cau…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1

Diversity or Precision? A Deep Dive into Next Token Prediction

2025-12-28 · Haoyuan Wu, Hai Wang, Jiajia Wu, Jinxiang Ou 외 arxiv

Recent advancements have shown that reinforcement learning (RL) can substantially improve the reasoning abilities of large language models (LLMs). The effectiveness of such RL training, however, depends critically on the…

Reinforcement Learning

Conditional Attribute Estimation with Autoregressive Sequence Models

2026-05-13 · Erica Stutz, Giacomo Marino, Daniella Meeker, Qiao Liu 외 arxiv

Generative models are often trained with a next-token prediction objective, yet many downstream applications require the ability to estimate or control sequence-level properties. Next-token prediction can lead to overfit…