paper-with-me

홈 › Papers

Estimating the Effective Rank of Vision Transformers via Low-Rank Factorization

2025-11-30 · Liyu Zerihun arxiv

Deep networks are heavily over-parameterized, yet their learned representations often admit low-rank structure. We introduce a framework for estimating a model's intrinsic dimensionality by treating learned representations as projections onto a low-rank subspace of the model's full capacity. Our approach: train a full-rank teacher, factorize its weights at multiple ranks, and train each factorized student via distillation to measure performance as a function of rank. We define effective rank as a region, not a point: the smallest contiguous set of ranks for which the student reaches 85-95% of teacher accuracy. To stabilize estimates, we fit accuracy vs. rank with a monotone PCHIP interpolant and identify crossings of the normalized curve. We also define the effective knee as the rank maximizing perpendicular distance between the smoothed accuracy curve and its endpoint secant; an intrinsic indicator of where marginal gains concentrate. On ViT-B/32 fine-tuned on CIFAR-100 (one seed, due to compute constraints), factorizing linear blocks and training with distillation yields an effective-rank region of approximately [16, 34] and an effective knee at r* ~ 31. At rank 32, the student attains 69.46% top-1 accuracy vs. 73.35% for the teacher (~94.7% of baseline) while achieving substantial parameter compression. We provide a framework to estimate effective-rank regions and knees across architectures and datasets, offering a practical tool for characterizing the intrinsic dimensionality of deep models.

📄 PDF Abstract BibTeX arXiv:2512.00792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Effective Fine-Tuning of Vision Transformers with Low-Rank Adaptation for Privacy-Preserving Image Classification

2025-07-16 · Haiwei Lin, Shoko Imaizumi, Hitoshi Kiya arxiv

We propose a low-rank adaptation method for training privacy-preserving vision transformer (ViT) models that efficiently freezes pre-trained ViT model weights. In the proposed method, trainable rank decomposition matrice…

Image Classification

Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility

2025-10-02 · Annan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang 외 arxiv

Transformers are widely used across data modalities, and yet the principles distilled from text models often transfer imperfectly to models trained to other modalities. In this paper, we analyze Transformers through the …

Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank Collapse

2022-06-07 · Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto 외

Transformers have achieved remarkable success in several domains, ranging from natural language processing to computer vision. Nevertheless, it has been recently shown that stacking self-attention layers - the distinctiv…

Collaborative Low-Rank Adaptation for Pre-Trained Vision Transformers

2025-12-31 · Zheng Liu, Jinchao Zhu, Gao Huang arxiv

Low-rank adaptation (LoRA) has achieved remarkable success in fine-tuning pre-trained vision transformers for various downstream tasks. Existing studies mainly focus on exploring more parameter-efficient strategies or mo…

Representation Learning

Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training

2025-11-06 · Ipsita Ghosh, Ethan Nguyen, Christian Kümmerle arxiv

Parameter-efficient training based on low-rank optimization has become a highly successful tool for fine-tuning large deep learning models. However, these methods often fail for low-rank pre-training, where simultaneousl…