paper-with-me

홈 › Papers

What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility

2026-06-19 · Yuren Hao, Xiang Wan, ChengXiang Zhai arxiv

Despite the success of deep learning, training deep networks in biologically plausible and hardware-efficient ways remains an open challenge. Feedback alignment (FA) methods address this by replacing backpropagation's symmetric backward weights with fixed random matrices, but their effectiveness depends critically on whether they can be accurately evaluated. The standard evaluation relies on two quantities: task accuracy and cosine similarity between the method's credit signal and the backpropagation gradient. We show that this reporting pair is insufficient by identifying two independent failure modes, both silent under current reporting: (1) measurement degeneracy, where the BP reference gradient collapses to the numerical floor in terminal-LayerNorm residual architectures, rendering cosine uninterpretable; and (2) aggregation collapse, where the aggregate cosine masks layerwise heterogeneity that concentrates credit at one end of the network. To address these limitations, we propose a diagnostic evaluation protocol based on three checks -- scale stability, reference validity, and depth utility -- together with per-layer rather than aggregate cosine reporting. Across multiple architectures and methods, the standard reporting pair gives no signal of failure in any audited case, while our protocol identifies all failures with wide calibration margins. The two failure modes are causally independent: a per-block scale penalty alleviates Mode 1 (residual scale explosion driving reference collapse) without affecting Mode 2 (cosine ranking that contradicts every functional metric we measured). Identifying these silent failures prevents researchers from building on non-functional credit assignment and provides actionable guidance for developing FA methods that genuinely train deep layers.

📄 PDF Abstract BibTeX arXiv:2606.21126

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks

2025-12-21 · Yogeswar Reddy Thota arxiv

Modern deep residual networks perform substantial redundant computation by evaluating all residual blocks for every input, even when identity mappings suffice. We introduce CosineGate, an end-to-end differentiable archit…

What Is Missing: Interpretable Ratings for Large Language Model Outputs

2026-02-17 · Nicholas Stranges, Yimin Yang arxiv

Current Large Language Model (LLM) preference learning methods such as Proximal Policy Optimization and Direct Preference Optimization learn from direct rankings or numerical ratings of model outputs, these rankings are …

Recoverable but Not Stationary:Local Linear Structures in Weights and Activations

2026-06-09 · Irina Piontkovskaia, Sergey Nikolenko arxiv

Task vectors, LoRA, activation steering, and random search around pretrained weights all suggest that learned behaviour can be controlled by linear directions. We ask which linear structures actually exist and on what sc…

Secure Cluster-Based Hierarchical Federated Learning in Vehicular Networks

2025-05-02 · M. Saeid HaghighiFard, Sinem Coleri

Hierarchical Federated Learning (HFL) has recently emerged as a promising solution for intelligent decision-making in vehicular networks, helping to address challenges such as limited communication resources, high vehicl…

Anomaly DetectionFederated Learning

Rethinking the Harmonic Loss via Non-Euclidean Distance Layers

2026-03-10 · Maxwell Miller-Golub, Collin Coil, Kamil Faber, Marcin Pietron 외 arxiv

Cross-entropy loss has long been the standard choice for training deep neural networks, yet it suffers from interpretability limitations, unbounded weight growth, and inefficiencies that can contribute to costly training…

Computational Efficiency