paper-with-me

홈 › Papers

SentenceVAE: Enable Next-sentence Prediction for Large Language Models with Faster Speed, Higher Accuracy and Longer Context

2024-08-01 · Hongjun An, Yifan Chen, Zhe Sun, Xuelong Li

Current large language models (LLMs) primarily utilize next-token prediction method for inference, which significantly impedes their processing speed. In this paper, we introduce a novel inference methodology termed next-sentence prediction, aiming at enhancing the inference efficiency of LLMs. We present Sentence Variational Autoencoder (SentenceVAE), which includes a Sentence Encoder to compress multiple tokens in a sentence into a single token, and a Sentence Decoder to reconstruct it. By integrating SentenceVAE into the input and output layers of LLMs, we develop Sentence-level LLMs (SLLMs) that employ a sentence-by-sentence inference method. In addition, the SentenceVAE module of SLLMs can maintain the integrity of the original semantic content by segmenting the context into sentences, thereby improving accuracy while boosting inference speed. Moreover, compared to previous LLMs, SLLMs process fewer tokens over equivalent context length, significantly reducing memory demands for self-attention computation and facilitating the handling of longer context. Extensive experiments on Wanjuan dataset have revealed that the proposed method can accelerate inference speed by 204~365%, reduce perplexity (PPL) to 46~75% of its original metric, and decrease memory overhead by 86~91% for the equivalent context length, compared to previous token-by-token methods.

📄 PDF Abstract BibTeX arXiv:2408.00655

Code (1)

BestAnHongjun/SentenceVAE 공식 구현 pytorch

Tasks

DecoderSentence

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Toward Better Storylines with Sentence-Level Language Models

2020-05-11 · ACL 2020 6 · Daphne Ippolito, David Grangier, Douglas Eck, Chris Callison-Burch

We propose a sentence-level language model which selects the next sentence in a story from a finite set of fluent alternatives. Since it does not need to model fluency, the sentence-level language model can focus on long…

Language ModelingLanguage ModellingSentenceSentence Embeddings+1

Assisting Composition of Email Responses: a Topic Prediction Approach

2015-10-07 · Spandana Gella, Marc Dymetman, Jean Michel Renders, Sriram Venkatapathy

We propose an approach for helping agents compose email replies to customer requests. To enable that, we use LDA to extract latent topics from a collection of email exchanges. We then use these latent topics to label our…

Sentence

Contextual LSTM (CLSTM) models for Large scale NLP tasks

2016-02-19 · Shalini Ghosh, Oriol Vinyals, Brian Strope, Scott Roy 외

Documents exhibit sequential structure at multiple levels of abstraction (e.g., sentences, paragraphs, sections). These abstractions constitute a natural hierarchy for representing the context in which to infer the meani…

ArticlesParaphrase GenerationPredictionQuestion Answering+2

What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction

2025-08-11 · Charlie Wyatt, Aditya Joshi, Flora Salim arxiv

Transformer-based models primarily rely on Next Token Prediction (NTP), which predicts the next token in a sequence based on the preceding context. However, NTP's focus on single-token prediction often limits a model's a…

Sentence Embeddings for Russian NLU

2019-10-29 · Dmitry Popov, Alexander Pugachev, Polina Svyatokum, Elizaveta Svitanko 외

We investigate the performance of sentence embeddings models on several tasks for the Russian language. In our comparison, we include such tasks as multiple choice question answering, next sentence prediction, and paraph…

Multiple-choiceParaphrase IdentificationPredictionQuestion Answering+2