paper-with-me

홈 › Papers

On the Residual Scaling of Looped Transformers: Stability and Transferability

2026-06-16 · Shaowen Wang, Bingrui Li, Ge Zhang, Wenhao Huang, Shen Yan, Jian Li arxiv

Looped (weight-tied) Transformers apply a shared residual block $N$ times ($h \leftarrow h + \varepsilon\,f(h)$, same $f$ at each step), increasing effective depth without adding parameters. Prior depth-scaling analyses prescribe $\varepsilon = 1/\!\sqrt{L}$ for depth-$L$ residual networks. We show that this is insufficient for looped architectures: weight sharing makes residual updates correlated across iterations, requiring the stronger scaling $\varepsilon = 1/N$. For multi-layer blocks ($L$ unique layers looped $N$ times), we derive a factored parameterization $\varepsilon = λ/(N\!\sqrt{L})$ that separates the two sources of growth: $1/N$ controls the within-layer loop correlation, and $1/\!\sqrt{L}$ controls the across-layer variance. A key consequence is that the optimal learning rate depends only on the number of unique layers $L$, not on the loop count $N$, enabling direct hyperparameter transfer from small to large $N$ without retuning. Experiments on looped Transformers confirm that $1/N$ scaling improves trainability and yields better loss than $1/\!\sqrt{N}$ scaling across loop counts.

📄 PDF Abstract BibTeX arXiv:2606.18524

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeepLoop: Depth Scaling for Looped Transformers

2026-07-15 · Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu 외 hf

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling…

Simply Stabilizing the Loop via Fully Looped Transformer

2026-05-11 · Rao Fu, Zixuan Yang, Jiankun Zhang, Jing Ma 외 arxiv

Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading additional computation for improved perfo…

Parcae: Scaling Laws For Stable Looped Language Models

2026-04-14 · Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick, Daniel Y. Fu arxiv

Traditional fixed-depth architectures scale quality by increasing training FLOPs, typically through increased parameterization, at the expense of a higher memory footprint, or data. A potential alternative is looped arch…

Stability and Generalization in Looped Transformers

2026-04-16 · Asher Labovich arxiv

Looped transformers promise test-time compute scaling by spending more iterations on harder problems, but it remains unclear which architectural choices let them extrapolate to harder problems at test time rather than me…

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

2026-06-16 · Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus 외 arxiv

Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solut…