paper-with-me

홈 › Papers

CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation

2024-10-23 · Qinsi Wang, Saeed Vahidian, Hancheng Ye, Jianyang Gu, Jianyi Zhang, Yiran Chen

Large language models (LLMs) with billions of parameters have sparked a new wave of exciting AI applications. However, their high computational costs and memory demands during inference pose significant challenges. Adaptive sparse activation inference, which activates only a small number of neurons for each token, offers a novel way to accelerate model inference without degrading performance, showing great potential for resource-constrained hardware devices. Nevertheless, existing methods predict activated neurons based on individual tokens with additional MLP, which involve frequent changes in activation maps and resource calls, limiting the acceleration benefits of sparse activation. In this paper, we introduce CoreInfer, an MLP-free adaptive sparse activation inference method based on sentence-level prediction. Specifically, we propose the concept of sentence-wise core neurons, which refers to the subset of neurons most critical for a given sentence, and empirically demonstrate its effectiveness. To determine the core neurons, we explore the correlation between core neurons and the sentence's semantics. Remarkably, we discovered that core neurons exhibit both stability and similarity in relation to the sentence's semantics -- an insight overlooked by previous studies. Building on this finding, we further design two semantic-based methods for predicting core neurons to fit different input scenarios. In CoreInfer, the core neurons are determined during the pre-filling stage and fixed during the encoding stage, enabling zero-cost sparse inference. We evaluated the model generalization and task generalization of CoreInfer across various models and tasks. Notably, on an NVIDIA TITAN XP GPU, CoreInfer achieved a 10.33 times and 2.72 times speedup compared to the Huggingface implementation and PowerInfer, respectively.

📄 PDF Abstract BibTeX arXiv:2410.18311

Code (0)

등록된 구현이 없습니다.

Tasks

GPULanguage ModelingLanguage ModellingLarge Language ModelSentence

Similar Papers 제목 키워드 기반

Semantics-Driven Cloud-Edge Collaborative Inference

2023-09-27 · Yuche Gao, Beibei Zhang

With the proliferation of video data in smart city applications like intelligent transportation, efficient video analytics has become crucial but also challenging. This paper proposes a semantics-driven cloud-edge collab…

Collaborative InferenceLicense Plate Recognition

Accelerating Production LLMs with Combined Token/Embedding Speculators

2024-04-29 · Davis Wertheimer, Joshua Rosenkranz, Thomas Parnell, Sahil Suneja 외

This technical report describes the design and training of novel speculative decoding draft models, for accelerating the inference speeds of large language models in a production environment. By conditioning draft predic…

Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing

2022-04-21 · Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C. Castro 외

Multi-modal data abounds in biomedicine, such as radiology images and reports. Interpreting this data at scale is essential for improving clinical care and accelerating clinical research. Biomedical text with its complex…

Contrastive LearningLanguage ModelingLanguage ModellingMedical Image Classification+3

HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models

2025-09-28 · Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng arxiv

Speculative decoding has proven effective for accelerating inference in Large Language Models (LLMs), yet its extension to Vision-Language Models (VLMs) remains limited by the computational burden and semantic inconsiste…

HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding

2026-01-13 · Qitan Lv, Tianyu Liu, Wen Wu, Xuenan Xu 외 arxiv

Speculative decoding (SD) has emerged as a promising approach to accelerate LLM inference without sacrificing output quality. Existing SD methods tailored for video-LLMs primarily focus on pruning redundant visual tokens…