paper-with-me

홈 › Papers

Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction

2024-09-25 · Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, YIngyu Liang, Shafiq Joty

Large Language Models (LLMs) have demonstrated remarkable capabilities in handling long context inputs, but this comes at the cost of increased computational resources and latency. Our research introduces a novel approach for the long context bottleneck to accelerate LLM inference and reduce GPU memory consumption. Our research demonstrates that LLMs can identify relevant tokens in the early layers before generating answers to a query. Leveraging this insight, we propose an algorithm that uses early layers of an LLM as filters to select and compress input tokens, significantly reducing the context length for subsequent processing. Our method, GemFilter, demonstrates substantial improvements in both speed and memory efficiency compared to existing techniques, such as standard attention and SnapKV/H2O. Notably, it achieves a 2.4$\times$ speedup and 30\% reduction in GPU memory usage compared to SOTA methods. Evaluation on the Needle in a Haystack task shows that GemFilter significantly outperforms standard attention, SnapKV and demonstrates comparable performance on the LongBench challenge. GemFilter is simple, training-free, and broadly applicable across different LLMs. Crucially, it provides interpretability by allowing humans to inspect the selected input sequence. These findings not only offer practical benefits for LLM deployment, but also enhance our understanding of LLM internal mechanisms, paving the way for further optimizations in LLM design and inference. Our code is available at \url{https://github.com/SalesforceAIResearch/GemFilter}.

📄 PDF Abstract BibTeX arXiv:2409.17422

Code (1)

salesforceairesearch/gemfilter 공식 구현 pytorch

Tasks

GPUToken Reduction

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Discovering Hidden Gems in Model Repositories

2026-01-29 · Jonathan Kahana, Eliahu Horwitz, Yedid Hoshen arxiv

Public repositories host millions of fine-tuned models, yet community usage remains disproportionately concentrated on a small number of foundation checkpoints. We investigate whether this concentration reflects efficien…

t-gems: text-guided exit modules for decreasing clip image encoder

2026-05-17 · Alberto Presta, Grzegorz Stefanski, Michal Byra, Krzysztof Arendt arxiv

Multimodal deep neural networks enhance deep comprehension by integrating diverse data modalities. Data from different modalities are typically projected into a shared latent space for similarity computation, but this pr…

GEMSS: A Variational Bayesian Method for Discovering Multiple Sparse Solutions in Classification and Regression Problems

2026-02-09 · Kateřina Henclová, Václav Šmídl arxiv

High-dimensional, underdetermined and highly correlated systems are common in data science practice, especially when analyzing physical measurements. In such settings, feature selection poses a fundamental challenge beca…

Gemtelligence: Accelerating Gemstone classification with Deep Learning

2023-05-31 · Tommaso Bendinelli, Luca Biggio, Daniel Nyfeler, Abhigyan Ghosh 외

The value of luxury goods, particularly investment-grade gemstones, is greatly influenced by their origin and authenticity, sometimes resulting in differences worth millions of dollars. Traditionally, human experts have …

ClassificationDeep Learning

Gems: Group Emotion Profiling Through Multimodal Situational Understanding

2025-07-30 · Anubhav Kataria, Surbhi Madan, Shreya Ghosh, Tom Gedeon 외 arxiv

Understanding individual, group and event level emotions along with contextual information is crucial for analyzing a multi-person social situation. To achieve this, we frame emotion comprehension as the task of predicti…