paper-with-me

홈 › Papers

Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers

2026-05-07 · Pengqi Lu arxiv

Scaling Diffusion Transformers (DiTs) to hundreds of layers introduces a structural vulnerability: networks can enter a silent, mean-dominated collapse state that homogenizes token representations and suppresses centered variation. Through mechanistic auditing, we isolate the trigger event of this collapse as Mean Mode Screaming (MMS). MMS can occur even when training appears stable, with a mean-coherent backward shock on residual writers that opens deep residual branches and drives the network into a mean-dominated state. We show this behavior is driven by an exact decomposition of these gradients into mean-coherent and centered components, compounded by the structural suppression of attention-logit gradients through the null space of the Softmax Jacobian once values homogenize. To address this, we propose Mean-Variance Split (MV-Split) Residuals, which combine a separately gained centered residual update with a leaky trunk-mean replacement. On a 400-layer single-stream DiT, MV-Split prevents the divergent collapse that crashes the un-stabilized baseline; it tracks close to the baseline's pre-crash trajectory while remaining substantially better than token-isotropic gating methods such as LayerScale across the full schedule. Finally, we present a 1000-layer DiT as a scale-validation run at boundary scales, establishing that the architecture remains stably trainable at extreme depth.

📄 PDF Abstract BibTeX arXiv:2605.06169

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Weighted Elastic Net Penalized Mean-Variance Portfolio Design and Computation

2015-10-15

It is well known that the out-of-sample performance of Markowitz's mean-variance portfolio criterion can be negatively affected by estimation errors in the mean and covariance. In this paper we address the problem by reg…

The Effect of Heteroscedasticity on Regression Trees

2016-06-16 · Will Ruth, Thomas Loughin

Regression trees are becoming increasingly popular as omnibus predicting tools and as the basis of numerous modern statistical learning ensembles. Part of their popularity is their ability to create a regression predicti…

regression

From Pixels to Patches: Pooling Strategies for Earth Embeddings

2026-03-02 · Isaac Corley, Caleb Robinson, Inbal Becker-Reshef, Juan M. Lavista Ferres arxiv

Geospatial foundation models increasingly expose pixel-level embedding products that can be downloaded and reused without access to the underlying encoder. In this setting, downstream tasks with patch- or region-level la…

Detection of Children Abuse by Voice and Audio Classification by Short-Time Fourier Transform Machine Learning implemented on Nvidia Edge GPU device

2023-07-27 · Jiuqi Yan, Yingxian Chen, W. W. T. Fok

The safety of children in children home has become an increasing social concern, and the purpose of this experiment is to use machine learning applied to detect the scenarios of child abuse to increase the safety of chil…

Abuse DetectionAudio ClassificationGPUimage-classification+1

Thinning a Wishart Random Matrix

2025-02-14 · Ameer Dharamshi, Anna Neufeld, Lucy L. Gao, Daniela Witten 외

Recent work has explored data thinning, a generalization of sample splitting that involves decomposing a (possibly matrix-valued) random variable into independent components. In the special case of a $n \times p$ random …