paper-with-me

홈 › Papers

Temperature check: theory and practice for training models with softmax-cross-entropy losses

2020-10-14 · Atish Agarwala, Jeffrey Pennington, Yann Dauphin, Sam Schoenholz

The softmax function combined with a cross-entropy loss is a principled approach to modeling probability distributions that has become ubiquitous in deep learning. The softmax function is defined by a lone hyperparameter, the temperature, that is commonly set to one or regarded as a way to tune model confidence after training; however, less is known about how the temperature impacts training dynamics or generalization performance. In this work we develop a theory of early learning for models trained with softmax-cross-entropy loss and show that the learning dynamics depend crucially on the inverse-temperature $\beta$ as well as the magnitude of the logits at initialization, $||\beta{\bf z}||_{2}$. We follow up these analytic results with a large-scale empirical study of a variety of model architectures trained on CIFAR10, ImageNet, and IMDB sentiment analysis. We find that generalization performance depends strongly on the temperature, but only weakly on the initial logit magnitude. We provide evidence that the dependence of generalization on $\beta$ is not due to changes in model confidence, but is a dynamical phenomenon. It follows that the addition of $\beta$ as a tunable hyperparameter is key to maximizing model performance. Although we find the optimal $\beta$ to be sensitive to the architecture, our results suggest that tuning $\beta$ over the range $10^{-2}$ to $10^1$ improves performance over all architectures studied. We find that smaller $\beta$ may lead to better peak performance at the cost of learning stability.

📄 PDF Abstract BibTeX arXiv:2010.07344

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training

2026-04-05 · Charafeddine Mouzouni arxiv

We model Mixture-of-Experts (MoE) token routing as a congestion game with a single effective parameter, the congestion coefficient gamma_eff, that quantifies the balance-quality tradeoff. Tracking gamma_eff across traini…

Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions

2025-11-03 · Samet Demir, Zafer Dogan arxiv

Pretrained Transformers can perform in-context learning (ICL) from a few demonstrations, but this ability can fail sharply when the test distribution differs from pretraining, a common deployment setting. We study attent…

On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

2017-04-03 · Bolin Gao, Lacra Pavel

In this paper, we utilize results from convex analysis and monotone operator theory to derive additional properties of the softmax function that have not yet been covered in the existing literature. In particular, we sho…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Cross-Entropy Is Load-Bearing: A Pre-Registered Scope Test of the K-Way Energy Probe on Bidirectional Predictive Coding

2026-04-23 · Jon-Paul Cacioli arxiv

Cacioli (2026) showed that the K-way energy probe on standard discriminative predictive coding networks reduces approximately to a monotone function of the log-softmax margin. The reduction rests on five assumptions, inc…

Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium

2025-08-28 · Christopher R. Lee-Jenkins arxiv

Decoding in large language models is often described as scoring tokens and normalizing with softmax. We give a minimal, self-contained account of this step as a constrained variational principle on the probability simple…