paper-with-me

홈 › Papers

Search for Efficient Large Language Models

2024-09-25 · Xuan Shen, Pu Zhao, Yifan Gong, Zhenglun Kong, Zheng Zhan, Yushu Wu, Ming Lin, Chao Wu, Xue Lin, Yanzhi Wang

Large Language Models (LLMs) have long held sway in the realms of artificial intelligence research. Numerous efficient techniques, including weight pruning, quantization, and distillation, have been embraced to compress LLMs, targeting memory reduction and inference acceleration, which underscore the redundancy in LLMs. However, most model compression techniques concentrate on weight optimization, overlooking the exploration of optimal architectures. Besides, traditional architecture search methods, limited by the elevated complexity with extensive parameters, struggle to demonstrate their effectiveness on LLMs. In this paper, we propose a training-free architecture search framework to identify optimal subnets that preserve the fundamental strengths of the original LLMs while achieving inference acceleration. Furthermore, after generating subnets that inherit specific weights from the original LLMs, we introduce a reformation algorithm that utilizes the omitted weights to rectify the inherited weights with a small amount of calibration data. Compared with SOTA training-free structured pruning works that can generate smaller networks, our method demonstrates superior performance across standard benchmarks. Furthermore, our generated subnets can directly reduce the usage of GPU memory and achieve inference acceleration. Code: https://github.com/shawnricecake/search-llm

📄 PDF Abstract BibTeX arXiv:2409.17372

Code (1)

shawnricecake/search-llm 공식 구현 pytorch

Tasks

GPUModel CompressionQuantization

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Enhancing Cloud-Based Large Language Model Processing with Elasticsearch and Transformer Models

2024-02-24 · Chunhe Ni, Jiang Wu, Hongbo Wang, Wenran Lu 외

Large Language Models (LLMs) are a class of generative AI models built using the Transformer network, capable of leveraging vast datasets to identify, summarize, translate, predict, and generate language. LLMs promise to…

Language ModelingLanguage ModellingLarge Language Model

Lost in Translation: Large Language Models in Non-English Content Analysis

2023-06-12 · Gabriel Nicholas, Aliya Bhatia

In recent years, large language models (e.g., Open AI's GPT-4, Meta's LLaMa, Google's PaLM) have become the dominant approach for building AI systems to analyze and generate language online. However, the automated system…

Translation

A Survey of GPT-3 Family Large Language Models Including ChatGPT and GPT-4

2023-10-04 · Katikapalli Subramanyam Kalyan

Large language models (LLMs) are a special class of pretrained language models obtained by scaling model size, pretraining corpus and computation. LLMs, because of their large size and pretraining on large volumes of tex…

Data AugmentationSelf-Supervised LearningSurveyTransfer Learning

Algorithmic Ghost in the Research Shell: Large Language Models and Academic Knowledge Creation in Management Research

2023-03-10 · Nigel Williams, Stanislav Ivanov, Dimitrios Buhalis

The paper looks at the role of large language models in academic knowledge creation based on a scoping review (2018 to January 2023) of how researchers have previously used the language model GPT to assist in the perform…

ArticlesLanguage ModelingLanguage ModellingManagement

Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions

2025-01-12 · Aidan Hogan, Xin Luna Dong, Denny Vrandečić, Gerhard Weikum

Much has been discussed about how Large Language Models, Knowledge Graphs and Search Engines can be combined in a synergistic manner. A dimension largely absent from current academic discourse is the user perspective. In…

Knowledge Graphs