paper-with-me

홈 › Papers

The Unreasonable Ineffectiveness of the Deeper Layers

2024-03-26 · Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts

We empirically study a simple layer-pruning strategy for popular families of open-weight pretrained LLMs, finding minimal degradation of performance on different question-answering benchmarks until after a large fraction (up to half) of the layers are removed. To prune these models, we identify the optimal block of layers to prune by considering similarity across layers; then, to "heal" the damage, we perform a small amount of finetuning. In particular, we use parameter-efficient finetuning (PEFT) methods, specifically quantization and Low Rank Adapters (QLoRA), such that each of our experiments can be performed on a single A100 GPU. From a practical perspective, these results suggest that layer pruning methods can complement other PEFT strategies to further reduce computational resources of finetuning on the one hand, and can improve the memory and latency of inference on the other hand. From a scientific perspective, the robustness of these LLMs to the deletion of layers implies either that current pretraining methods are not properly leveraging the parameters in the deeper layers of the network or that the shallow layers play a critical role in storing knowledge.

📄 PDF Abstract BibTeX arXiv:2403.17887

Code (3)

RitAreaSciencePark/ZigZagLLMs pytorch
arcee-ai/PruneMe pytorch
mts-ai/replaceme pytorch

Tasks

GPUQuantizationQuestion Answering

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

The Curse of Depth in Large Language Models

2025-02-09 · Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin 외

In this paper, we introduce the Curse of Depth, a concept that highlights, explains, and addresses the recent observation in modern Large Language Models(LLMs) where nearly half of the layers are less effective than expe…

Deeplite Neutrino: An End-to-End Framework for Constrained Deep Learning Model Optimization

2021-01-11 · Anush Sankaran, Olivier Mastropietro, Ehsan Saboori, Yasser Idris 외

Designing deep learning-based solutions is becoming a race for training deeper models with a greater number of layers. While a large-size deeper model could provide competitive accuracy, it creates a lot of logistical ch…

Deep LearningModel Optimization

The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization

2024-08-29 · Luka Borec, Philipp Sadler, David Schlangen

This work analyses the text memorization behavior of large language models (LLMs) when subjected to nucleus sampling. Stochastic decoding methods like nucleus sampling are typically applied to overcome issues such as mon…

DiagnosticMemorizationText Generation

CompactifAI: Extreme Compression of Large Language Models using Quantum-Inspired Tensor Networks

2024-01-25 · Andrei Tomut, Saeed S. Jahromi, Abhijoy Sarkar, Uygar Kurt 외

Large Language Models (LLMs) such as ChatGPT and LlaMA are advancing rapidly in generative Artificial Intelligence (AI), but their immense size poses significant challenges, such as huge training and inference costs, sub…

Model CompressionQuantizationTensor Networks

Knowledge will Propel Machine Understanding of Content: Extrapolating from Current Examples

2017-07-14 · Amit Sheth, Sujan Perera, Sanjaya Wijeratne, Krishnaprasad Thirunarayan

Machine Learning has been a big success story during the AI resurgence. One particular stand out success relates to learning from a massive amount of data. In spite of early assertions of the unreasonable effectiveness o…