paper-with-me

홈 › Papers

Multi-Gate Residuals

2026-05-22 · Zhizhan Zheng, Feiyun Zhang, Shuchun Liu, Tian Xia, Xi Liu, Dasheng Hu, Hongquan Zhou arxiv

While Attention Residuals has shown some effectiveness in addressing the widespread issue of unbounded activation growth across deep residual layers, it inevitably incurs significant communication overhead. To circumvent this bottleneck, we propose Multi-Gate Residuals (MGR), which stabilizes activation scales without additional communication burden. It utilizes a straightforward scoring and gating mechanism to maintain multi-stream context, coupled with Attention Pooling to extract hidden states from the stream states. Empirical experiments demonstrate that MGR is practical for large-scale training and deployment, offering tangible performance improvements over existing architectures.

📄 PDF Abstract BibTeX arXiv:2605.23259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Review Residuals: Update-Conditioned Residual Gating for Transformers

2026-06-30 · Kyle Kramer arxiv

Residual connections add every sublayer's proposed update with a fixed coefficient of one; the network never evaluates whether an update is reliable before committing it. Drawing on the human-factors principle of indepen…

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

2026-07-29 · Haotian Mo, Jie Liu, Siqi Shen, Songzhu Mei 외 arxiv

Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions a…

Audio Deepfake Detection

Self-Navigated Residual Mamba for Universal Industrial Anomaly Detection

2025-08-03 · Hanxi Li, Jingqi Wu, Lin Yuanbo Wu, Mingliang Li 외 arxiv

In this paper, we propose Self-Navigated Residual Mamba (SNARM), a novel framework for universal industrial anomaly detection that leverages ``self-referential learning'' within test images to enhance anomaly discriminat…

Ensemble LearningAnomaly Detection

Multi-Head Attention Residuals

2026-07-22 · Cheng Luo, Zefan Cai, Junjie Hu hf

Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by letting each sublayer attend, through a learne…

Self and mutually exciting point process embedding flexible residuals and intensity with discretely Markovian dynamics

2024-01-25 · Kyungsub Lee

This work introduces a self and mutually exciting point process that embeds flexible residuals and intensity with discretely Markovian dynamics. By allowing the integration of diverse residual distributions, this model s…

Time Series