paper-with-me

Papers

Spectral Alignment as Predictor of Loss Explosion in Neural Network Training

2025-10-05 · Haiquan Qiu, You Wu, Yingjie Tan, Yaqing Wang, Quanming Yao arxiv

Loss explosions in training deep neural networks can nullify multi-million dollar training runs. Conventional monitoring metrics like weight and gradient norms are often lagging and ambiguous predictors, as their values vary dramatically across different models and even between layers of the same model, making it difficult to establish a unified standard for detecting impending failure. We introduce Spectral Alignment (SA), a novel, theoretically-grounded metric that monitors the distributional alignment between layer inputs and the principal singular vectors of weight matrices. We show that a collapse in the sign diversity of this alignment is a powerful early predictor of representational collapse and training divergence. Empirical results on language models demonstrate that monitoring the SA distribution provides a significantly earlier and clearer warning of loss explosions than traditional scalar metrics. SA's low computational overhead makes it a practical tool for safeguarding model training.

📄 PDF Abstract BibTeX arXiv:2510.04202

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment

2026-06-30 · Jason R. Brown, Patrick Leask, Lev McKinney arxiv

Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned behaviour on unrelated prompts. Previous wo…

Hyperspectral Anomaly Change Detection Based on Auto-encoder

2020-10-27 · Meiqi Hu, Chen Wu, Liangpei Zhang, Bo Du

With the hyperspectral imaging technology, hyperspectral data provides abundant spectral information and plays a more important role in geological survey, vegetation analysis and military reconnaissance. Different from n…

Change Detection

A Universal Trade-off Between the Model Size, Test Loss, and Training Loss of Linear Predictors

2022-07-23 · Nikhil Ghosh, Mikhail Belkin

In this work we establish an algorithm and distribution independent non-asymptotic trade-off between the model size, excess test loss, and training loss of linear predictors. Specifically, we show that models that perfor…

Lossless Compression of Mosaic Images with Convolutional Neural Network Prediction

2020-01-28 · Seyed Mehdi Ayyoubzadeh, Xiaolin Wu

We present a CNN-based predictive lossless compression scheme for raw color mosaic images of digital cameras. This specialized application problem was previously understudied but it is now becoming increasingly important…

DeblurringImage Restoration

Perceptual-Neural-Physical Sound Matching

2023-01-07 · Han Han, Vincent Lostanlen, Mathieu Lagrange

Sound matching algorithms seek to approximate a target waveform by parametric audio synthesis. Deep neural networks have achieved promising results in matching sustained harmonic tones. However, the task is more challeng…

AttributeAudio SynthesisDecoder