paper-with-me

홈 › Papers

Block-wise Training of Residual Networks via the Minimizing Movement Scheme

2022-10-03 · Skander Karkar, Ibrahim Ayed, Emmanuel de Bézenac, Patrick Gallinari

End-to-end backpropagation has a few shortcomings: it requires loading the entire model during training, which can be impossible in constrained settings, and suffers from three locking problems (forward locking, update locking and backward locking), which prohibit training the layers in parallel. Solving layer-wise optimization problems can address these problems and has been used in on-device training of neural networks. We develop a layer-wise training method, particularly welladapted to ResNets, inspired by the minimizing movement scheme for gradient flows in distribution space. The method amounts to a kinetic energy regularization of each block that makes the blocks optimal transport maps and endows them with regularity. It works by alleviating the stagnation problem observed in layer-wise training, whereby greedily-trained early layers overfit and deeper layers stop increasing test accuracy after a certain depth. We show on classification tasks that the test accuracy of block-wise trained ResNets is improved when using our method, whether the blocks are trained sequentially or in parallel.

📄 PDF Abstract BibTeX arXiv:2210.00949

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Module-wise Training of Neural Networks via the Minimizing Movement Scheme

2023-09-29 · NeurIPS 2023 11

Greedy layer-wise or module-wise training of neural networks is compelling in constrained and on-device settings where memory is limited, as it circumvents a number of problems of end-to-end back-propagation. However, it…

Distillation Guided Residual Learning for Binary Convolutional Neural Networks

2020-07-10 · Jianming Ye, Shiliang Zhang, Jingdong Wang

It is challenging to bridge the performance gap between Binary CNN (BCNN) and Floating point CNN (FCNN). We observe that, this performance gap leads to substantial residuals between intermediate feature maps of BCNN and …

Industrial Anomaly Detection and Localization Using Weakly-Supervised Residual Transformers

2023-06-06 · Hanxi Li, Jingqi Wu, Deyin Liu, Lin Wu 외

Recent advancements in industrial anomaly detection (AD) have demonstrated that incorporating a small number of anomalous samples during training can significantly enhance accuracy. However, this improvement often comes …

Anomaly DetectionAnomaly LocalizationSupervised Anomaly DetectionUnsupervised Anomaly Detection

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

2026-05-19 · Han Guo, Jack Zhang, Arjun Menon, Driss Guessous 외 arxiv

Transformer training systems are built around dense linear algebra, yet a nontrivial fraction of end-to-end time is spent on surrounding memory-bound operators. Normalization, activations, residual updates, reductions, a…

MIDUS: Memory-Infused Depth Up-Scaling

2025-12-15 · Taero Kim, Hoyoon Byun, Youngjun Choi, Sungrae Park 외 arxiv

Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scaling (DUS) does so by duplicating Transformer blocks and inserting them into a pr…