paper-with-me

홈 › Papers

Depth-Staggered Fibonacci Spacing for Sparse Attention: Static Schedules Beat Learned Dilation and Extrapolate Where Dense Attention Fails

2026-06-26 · Chad A. Capps arxiv

We study sparse self-attention in which each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alpha that compresses or expands the spacing. Across 21 language models trained under one matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four ways of setting alpha across depth: fixed, per-layer learned, a static linear stagger, and a coprime (anti-gridding) reassignment of that stagger, together with a reach-matched power-of-2 control. Three results stand out. First, a static per-layer stagger improves perplexity over both fixed and learned alpha, and the gain is base-agnostic: applying the same stagger to a power-of-2 base lifts it above fixed Fibonacci and to parity with learned Fibonacci attention. Second, learning per layer is inert: it does not beat the static schedule and costs roughly five times the inference latency. Third, and most consequential, all sparse variants extrapolate to four times their training length with little or no degradation, whereas a recipe-matched dense baseline collapses (perplexity rises by 201% at 4x length); we attribute this to fixed-offset attention only ever querying relative positions seen during training. We also report two honest negatives: at training length the best sparse model has about 26% higher perplexity than the dense baseline, and the staggering gain is uniform across context positions rather than concentrated at long range.

📄 PDF Abstract BibTeX arXiv:2606.28560

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads

2024-06-27 · Ali Khaleghi Rahimian, Manish Kumar Govind, Subhajit Maity, Dominick Reilly 외

Transformer architectures such as Vision Transformers (ViT) have proven effective for solving visual perception tasks. However, they suffer from two major limitations; first, the quadratic complexity of self-attention li…

Diversityimage-classificationImage ClassificationRepresentation Learning+1

On Fibonacci Ensembles: An Alternative Approach to Ensemble Learning Inspired by the Timeless Architecture of the Golden Ratio

2025-12-25 · Ernest Fokoué arxiv

Nature rarely reveals her secrets bluntly, yet in the Fibonacci sequence she grants us a glimpse of her quiet architecture of growth, harmony, and recursive stability \citep{Koshy2001Fibonacci, Livio2002GoldenRatio}. Fro…

Ensemble Learning

Fibonacci-Net: A Lightweight CNN model for Automatic Brain Tumor Classification

2025-03-18 · Santanu Roy, Ashvath Suresh, Archit Gupta, Shubhi Tiwari 외

This research proposes a very lightweight model "Fibonacci-Net" along with a novel pooling technique, for automatic brain tumor classification from imbalanced Magnetic Resonance Imaging (MRI) datasets. Automatic brain tu…

Brain Tumor ClassificationSpecificity

Disk-stacking models are consistent with Fibonacci and non-Fibonacci structure in sunflowers

2024-07-08 · Jonathan Swinton

This paper investigates a model of plant organ placement motivated by the appearance of large Fibonacci numbers in phyllotaxis, and provides the first large-scale empirical validation of this model. Specifically it evalu…

Second Maximum of a Gaussian Random Field and Exact (t-)Spacing test

2024-06-26 · Jean-Marc Azaïs, Federico Dalmao, Yohann de Castro

In this article, we introduce the novel concept of the second maximum of a Gaussian random field on a Riemannian submanifold. This second maximum serves as a powerful tool for characterizing the distribution of the maxim…