paper-with-me

홈 › Papers

Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training

2026-02-21 · Seyed Morteza Emadi arxiv

Attention scores in transformers are bilinear forms $S_{ij} = x_i^\top M x_j / \sqrt{d_h}$ whose maximum magnitude governs overflow risk in low-precision training. We derive a \emph{rank-aware concentration inequality}: when the interaction matrix $M = W^Q W^{K\top}$ has rank $r \ll d$, tail probabilities for $\max_{i,j}|S_{ij}|$ decay as $\exp(-d^{2}α^{2}/(γr))$ rather than $\exp(-dα^{2})$, where $γ> 1$ is a typicality parameter. For transformer attention where $r = d_h$, this yields $8$--$28\times$ tighter concentration than rank-agnostic bounds in modern architectures. We apply this result to FP8 training, deriving \emph{geometry-aware scale factors} that provide principled overflow guarantees without observing activations. The method computes per-layer scales from the spectral norm $\|W^Q W^{K\top}\|_2$ via implicit power iteration, includes a grouped query attention formulation that avoids key expansion, and remains compatible with fused attention kernels. Across GPT-2 XL to Llama-2-70B, geometry-aware scaling eliminates overflows in transient scenarios where delayed scaling fails, while achieving comparable downstream MMLU accuracy.

📄 PDF Abstract BibTeX arXiv:2602.18851

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

2026-08-26 · Gerard Conangla Planes arxiv

Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. …

Perturbation Bounds for Low-Rank Inverse Approximations under Noise

2025-10-29 · Phuc Tran, Nisheeth K. Vishnoi arxiv

Low-rank pseudoinverses are widely used to approximate matrix inverses in scalable machine learning, optimization, and scientific computing. However, real-world matrices are often observed with noise, arising from sampli…

Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation

2025-09-19 · Jin Li, Zhebo Wang, Tianliang Lu, Mohan Li 외 arxiv

Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization techniques, suffer from high computation…

Text Generation

Adaptive Band Selection for Hyperspectral Classification with Spatially Disjoint Evaluation

2026-06-04 · Ikram El-Hajri, Ouassim Karrakchou, Alejandro Mousist arxiv

Hyperspectral band selection methods based on differentiable selectors can be sensitive to initialization and to extracting a final discrete subset, while prescribed band counts limit flexibility. We propose SGBR-HC (Spe…

LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift

2026-05-18 · Haozhe Si, Yuxuan Wan, Yuqing Wang, Minh Do 외 arxiv

Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling, and channel dimensionality. As a result, models trained under a fixe…

Representation Learning