paper-with-me

홈 › Papers

PocketLLM: Ultimate Compression of Large Language Models via Meta Networks

2025-11-19 · Ye Tian, Chengcheng Wang, Jing Han, Yehui Tang, Kai Han arxiv

As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs without sacrificing accuracy. In this paper, we introduce PocketLLM, a novel approach to compress LLMs in a latent space via meta-networks. A simple encoder network is proposed to project the weights of LLMs into discrete latent vectors, which are then represented using a compact codebook. A lightweight decoder network is employed to map the codebook's representative vectors back to the original weight space. This method allows for significant compression of the large weights in LLMs, consisting solely of a small decoder, a concise codebook, and an index. Extensive experiments show that PocketLLM achieves superior performance even at significantly high compression ratios, e.g., compressing Llama 2-7B by 10x with a negligible drop in accuracy.

📄 PDF Abstract BibTeX arXiv:2511.17637

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lossless Compression for LLM Tensor Incremental Snapshots

2025-05-14 · Daniel Waddington, Cornel Constantinescu

During the training of Large Language Models (LLMs), tensor data is periodically "checkpointed" to persistent storage to allow recovery of work done in the event of failure. The volume of data that must be copied during …

CPU

PocketLLM: Enabling On-Device Fine-Tuning for Personalized LLMs

2024-07-01 · Dan Peng, Zhihui Fu, Jun Wang

Recent advancements in large language models (LLMs) have indeed showcased their impressive capabilities. On mobile devices, the wealth of valuable, non-public data generated daily holds great promise for locally fine-tun…

I am a Strange Dataset: Metalinguistic Tests for Language Models

2024-01-10 · Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts 외

Statements involving metalinguistic self-reference ("This paper has six sections.") are prevalent in many domains. Can current large language models (LLMs) handle such language? In this paper, we present "I am a Strange …

Sentence

Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across Domains

2020-12-02 · ACL 2021 5 · Haojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang 외

Pre-trained language models have been applied to various NLP tasks with considerable performance gains. However, the large model sizes, together with the long inference time, limit the deployment of such models in real-t…

Knowledge DistillationLanguage ModelingLanguage ModellingMeta-Learning+2

Digital Modeling on Large Kernel Metamaterial Neural Network

2023-07-21 · Quan Liu, Hanyu Zheng, Brandon T. Swartz, Ho Hin Lee 외

Deep neural networks (DNNs) utilized recently are physically deployed with computational units (e.g., CPUs and GPUs). Such a design might lead to a heavy computational burden, significant latency, and intensive power con…

Edge-computing