paper-with-me

Papers

CORP: Closed-Form One-shot Representation-Preserving Structured Pruning for Transformers

2026-02-05 · Boxiang Zhang, Baijian Yang arxiv

Transformers achieve strong accuracy but incur high compute and memory cost. Structured pruning reduces inference cost, but most methods rely on retraining or multi-stage optimization, which limits post-training deployment. We propose CORP, a closed-form one-shot structured pruning method that removes MLP dimensions and attention substructures using only unlabeled calibration data without gradients or fine-tuning. CORP formulates structured pruning as a representation recovery problem. It models removed components as affine functions of retained components and derives closed-form ridge regression solutions that fold compensation into model weights. This minimizes a layer-local affine/logit reconstruction objective under the calibration distribution. Experiments on ImageNet with DeiT reveal strong redundancy in both MLP and attention representations. With CORP, models retain high accuracy under aggressive sparsity. On DeiT-Huge, CORP achieves 83.27% Top-1 accuracy after pruning 50\% of both MLP and attention structures.

📄 PDF Abstract BibTeX arXiv:2602.05243

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

2026-05-16 · Yujie Lin, Chengyi Yang, Zhishang Xiang, Yiping Song 외 arxiv

Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web corpora, raising concerns for privacy and safety. Existing machine unle…

Expected utility theory on mixture spaces without the completeness axiom

2021-02-13 · David McCarthy, Kalle Mikkola, Teruji Thomas

A mixture preorder is a preorder on a mixture space (such as a convex set) that is compatible with the mixing operation. In decision theoretic terms, it satisfies the central expected utility axiom of strong independence…

Open-Ended Question Answering

SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS

2026-01-23 · Ayush Pratap Singh, Harshit Singh, Nityanand Mathur, Akshat Mandloi 외 arxiv

Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepresentation in predominantly English train…

knowledge editing

Neural Cages for Detail-Preserving 3D Deformations

2019-12-13 · CVPR 2020 6 · Wang Yifan, Noam Aigerman, Vladimir G. Kim, Siddhartha Chaudhuri 외

We propose a novel learnable representation for detail-preserving shape deformation. The goal of our method is to warp a source shape to match the general structure of a target shape, while preserving the surface details…

Prototypical Model with Novel Information-theoretic Loss Function for Generalized Zero Shot Learning

2021-12-06 · Chunlin Ji, Hanchu Shen, Zhan Xiong, Feng Chen 외

Generalized zero shot learning (GZSL) is still a technical challenge of deep learning as it has to recognize both source and target classes without data from target classes. To preserve the semantic relation between sour…

Generalized Zero-Shot LearningRelationTransfer LearningZero-Shot Learning