paper-with-me

홈 › Papers

LiveMind: Low-latency Large Language Models with Simultaneous Inference

2024-06-20 · Chuangtao Chen, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li

In this paper, we introduce LiveMind, a novel low-latency inference framework for large language model (LLM) inference which enables LLMs to perform inferences with incomplete user input. By reallocating computational processes to the input phase, a substantial reduction in latency is achieved, thereby significantly enhancing the interactive experience for users of LLMs. The framework adeptly manages the visibility of the streaming input to the model, allowing it to infer from incomplete user input or await additional content. Compared with traditional inference methods on complete user input, our approach demonstrates an average reduction in response latency of 84.0% on the MMLU dataset and 71.6% on the MMLU-Pro dataset, while maintaining comparable accuracy. Additionally, our framework facilitates collaborative inference and output across different models. By employing an large LLM for inference and a small LLM for output, we achieve an average 37% reduction in response latency, alongside a 4.30% improvement in accuracy on the MMLU-Pro dataset compared with the baseline. The proposed LiveMind framework advances the field of human-AI interaction by enabling more responsive and efficient communication between users and AI systems.

📄 PDF Abstract BibTeX arXiv:2406.14319

Code (1)

chuangtaochen-tum/livemind 공식 구현

Tasks

Collaborative InferenceLanguage ModelingLanguage ModellingLarge Language ModelMMLU

Similar Papers 제목 키워드 기반

Conversational SimulMT: Efficient Simultaneous Translation with Large Language Models

2024-02-16 · Minghan Wang, Thuy-Trang Vu, Yuxia Wang, Ehsan Shareghi 외

Simultaneous machine translation (SimulMT) presents a challenging trade-off between translation quality and latency. Recent studies have shown that LLMs can achieve good performance in SimulMT tasks. However, this often …

Machine TranslationTranslation

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation

2025-04-22 · Keqi Deng, Wenxi Chen, Xie Chen, Philip C. Woodland

Simultaneous speech translation (SST) outputs translations in parallel with streaming speech input, balancing translation quality and latency. While large language models (LLMs) have been extended to handle the speech mo…

Simultaneous Speech-to-Speech TranslationSpeech-to-Speech TranslationTranslation

FASST: Fast LLM-based Simultaneous Speech Translation

2024-08-18 · Siqi Ouyang, Xi Xu, Chinmay Dandekar, Lei LI

Simultaneous speech translation (SST) takes streaming speech input and generates text translation on the fly. Existing methods either have high latency due to recomputation of input representations, or fall behind of off…

Language ModelingLanguage ModellingLarge Language ModelTranslation

Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time Scaling

2025-11-12 · Shiyu Ji, Yixuan Wang, Yijun Liu, Qingfu Zhu 외 arxiv

Test-time scaling improves the inference performance of Large Language Models (LLMs) but also incurs substantial computational costs. Although recent studies have reduced token consumption through dynamic self-consistenc…

PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity Recognition

2024-02-07 · Jinghui Lu, Ziwei Yang, Yanjie Wang, Xuejing Liu 외

In this study, we aim to reduce generation latency for Named Entity Recognition (NER) with Large Language Models (LLMs). The main cause of high latency in LLMs is the sequential decoding process, which autoregressively g…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER