paper-with-me

홈 › Papers

Suppressing Final Layer Hidden State Jumps in Transformer Pretraining

2026-01-26 · Keigo Shibata, Kazuki Yano, Ryosuke Takahashi, Jaesung Lee, Wataru Ikeda, Jun Suzuki arxiv

This paper discusses the internal behavior of Transformer language models. Many recent pre-trained models have been reported to exhibit only slight changes in the angular distance between the input and output hidden state vectors in the middle Transformer layers, despite a disproportionately large ``jump'' in the angular distance occurring in or around the final Transformer layer. To characterize this, we first introduce a quantitative metric for the jump strength around the final layer, and then demonstrate its prevalence across many open-weight models, as well as its amplification throughout pre-training. Assuming such jumps indicate an undesirable property, we propose the jump-suppressing regularizer (JREG) which penalizes this jump during pre-training, thereby encouraging more balanced capability usage across the middle layers. Empirical evaluations of three model sizes of Llama-based models, trained with the proposed JREG method, reveal improved task performance compared to the baseline without altering the model architecture.

📄 PDF Abstract BibTeX arXiv:2601.18302

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Multi-Stage Temporal Convolutional Network for Volleyball Jumps Classification Using a Waist-Mounted IMU

2023-10-19 · Meng Shang, Camilla De Bleecker, Jos Vanrenterghem, Roel De Ridder 외

Monitoring the number of jumps for volleyball players during training or a match can be crucial to prevent injuries, yet the measurement requires considerable workload and cost using traditional methods such as video ana…

Discontinuity-preserving Normal Integration with Auxiliary Edges

2024-04-04 · CVPR 2024 1 · Hyomin Kim, Yucheol Jung, Seungyong Lee

Many surface reconstruction methods incorporate normal integration, which is a process to obtain a depth map from surface gradients. In this process, the input may represent a surface with discontinuities, e.g., due to s…

Surface Reconstruction

One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels

2025-04-18 · Amrit Diggavi Seshadri

To reduce the time and computational costs of inference of large language models, there has been interest in parameter-efficient low-rank early-exit casting of transformer hidden-representations to final-representations.…

All

Backdoor Defense via Suppressing Model Shortcuts

2022-11-02 · Sheng Yang, Yiming Li, Yong Jiang, Shu-Tao Xia

Recent studies have demonstrated that deep neural networks (DNNs) are vulnerable to backdoor attacks during the training process. Specifically, the adversaries intend to embed hidden backdoors in DNNs so that malicious m…

backdoor defensemodel

Making Sense of Hidden Layer Information in Deep Networks by Learning Hierarchical Targets

2015-05-03 · Abhinav Tushar

This paper proposes an architecture for deep neural networks with hidden layer branches that learn targets of lower hierarchy than final layer targets. The branches provide a channel for enforcing useful information in h…

General Classificationtext-classificationText Classification