paper-with-me

홈 › Papers

Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training

2026-05-13 · Namrata Shivagunde, Vijeta Deshpande, Sherin Muckatira, Anna Rumshisky arxiv

Pre-training large language models is dominated by the memory cost of storing full-rank weights, gradients, and optimizer states. Low-rank pre-training has emerged to address this, and the space of methods has grown rapidly. A central question remains open: do low-rank methods produce models that generalize comparably to full-rank training, or does the rank constraint fundamentally alter the solutions reached? Existing comparisons rely almost entirely on validation perplexity from single-seed runs, often carried forward from prior literature. Yet perplexity is a poor proxy for solution quality; two methods can match on perplexity while converging to different loss landscape regions and internal representations. We close this gap by characterizing the solutions found by five low-rank pre-training methods, GaLore and Fira (memory-efficient optimizers), CoLA and SLTrain (architecture reparameterizations), and ReLoRA (adapter-style updates with periodic resets), against full-rank training at three model scales (60M, 130M, 350M). We evaluate each along 16 metrics across four dimensions: 1-D loss landscape along random/top-K PCA directions, 1-D interpolation between checkpoints, spectral structure of the weights and learned updates, and activation similarity to full-rank training. We show that low-rank methods are not equivalent to full-rank training, nor to one another, even when validation perplexity is close. Full-rank training settles into a sharper basin than low-rank methods along random directions, while the reverse holds for the top-1 PCA direction. Each method converges to a geometrically distinct basin. Low-rank activations diverge from full-rank in later layers as training progresses, with GaLore tracking full-rank most closely. Further, validation perplexity does not translate to downstream performance at every scale. Adding geometric and spectral metrics improves the prediction.

📄 PDF Abstract BibTeX arXiv:2605.13652

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction

2026-04-01 · Björn Roman Kohlberger arxiv

The memory wall remains the primary bottleneck for training large language models on consumer hardware. We introduce Spectral Compact Training (SCT), a method that replaces dense weight matrices with permanent truncated …

Low-Rank Decay for Grokking in Scale-Invariant Transformers: A Spectral-Geometric View

2026-06-03 · Mingyu Li arxiv

Modern Transformer architectures frequently employ normalization mechanisms such as RMSNorm and Query-Key Normalization, making parts of the model approximately scale-invariant with respect to weight magnitudes. In this …

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives

2026-05-01 · Mohammed Abu Baker, Luca Baroni, Dan Wilhelm arxiv

Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study these risks, researchers develop model organisms: models finetuned to exhibit specifi…

Spectral Geometric Verification: Re-Ranking Point Cloud Retrieval for Metric Localization

2022-10-10 · Kavisha Vidanapathirana, Peyman Moghadam, Sridha Sridharan, Clinton Fookes

In large-scale metric localization, an incorrect result during retrieval will lead to an incorrect pose estimate or loop closure. Re-ranking methods propose to take into account all the top retrieval candidates and re-or…

Point Cloud RegistrationPoint Cloud RetrievalPose EstimationRe-Ranking+1

Beyond Rigid Geometries: The Spline-Pullback Metric for Universal Diffeomorphic SPD Representation Learning

2026-05-06 · Tushar Das, Subrata Dutta, Sarmistha Neogy, Koushlendra Kumar Singh arxiv

The integration of Symmetric Positive Definite (SPD) matrices into deep learning has historically relied on fixed algebraic Riemannian metrics. Analogous to hand-crafted features in classical machine learning, these stat…

Representation Learning