paper-with-me

홈 › Papers

Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

2026-05-15 · Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy, Sunwoo Lee arxiv

Layer pruning removes entire Transformer decoder blocks from large language models, but introduces a mismatch between the hidden state received by the next surviving layer and the distribution it was trained to process, leading to significant performance degradation. We propose Ghosted Layers, a training-free recovery module that addresses this issue by solving a boundary activation alignment problem. Our method derives a closed-form optimal linear operator from a small calibration set to reconstruct the activation discrepancy introduced by the pruned layers. We show that this solution corresponds to the unconstrained optimum of the alignment objective, whereas existing methods are restricted to constrained solutions over limited operator subspaces. Experiments across multiple LLM backbones and pruning strategies demonstrate that our method consistently improves accuracy and perplexity over prior training-free baselines, while preserving the efficiency gains of layer pruning. Official code repository: https://github.com/daniel-eai/ghosted_layers_official_repository/.

📄 PDF Abstract BibTeX arXiv:2605.15491

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reflection Removal Using Ghosting Cues

2015-06-01 · CVPR 2015 6 · YiChang Shih, Dilip Krishnan, Fredo Durand, William T. Freeman

Photographs taken through glass windows often contain both the desired scene and undesired reflections. Separating the reflection and transmission layers is an important but ill-posed problem that has both aesthetic and …

Reflection Removal

Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment

2026-05-20 · Haozhe Jia, Pengyu Yin, Wenshuo Chen, Shaofeng Liang 외 arxiv

Physics-informed diffusion models typically enforce PDE constraints only on final outputs, leaving intermediate representations unconstrained and prone to shortcut learning under shifted boundary conditions. We introduce…

Recovering the Lowest Layer of Deep Networks with High Threshold Activations

2019-03-21 · ICLR 2019 5 · Surbhi Goel, Rina Panigrahy

Giving provable guarantees for learning neural networks is a core challenge of machine learning theory. Most prior work gives parameter recovery guarantees for one hidden layer networks, however, the networks used in pra…

BIG-bench Machine LearningLearning TheoryVocal Bursts Intensity Prediction

Steering Vision-Language Models with Joint Sparse Autoencoders

2026-06-24 · Huizhen Shu, Xuying Li, Hongxu Lin, Wenjie Sun 외 arxiv

Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering …

Gloss Alignment Using Word Embeddings

2023-08-08 · Harry Walsh, Ozge Mercanoglu Sincan, Ben Saunders, Richard Bowden

Capturing and annotating Sign language datasets is a time consuming and costly process. Current datasets are orders of magnitude too small to successfully train unconstrained \acf{slt} models. As a result, research has t…

Word AlignmentWord Embeddings