paper-with-me

Papers

Deep orthogonal linear networks are shallow

2020-11-27 · Pierre Ablin

We consider the problem of training a deep orthogonal linear network, which consists of a product of orthogonal matrices, with no non-linearity in-between. We show that training the weights with Riemannian gradient descent is equivalent to training the whole factorization by gradient descent. This means that there is no effect of overparametrization and implicit bias at all in this setting: training such a deep, overparametrized, network is perfectly equivalent to training a one-layer shallow network.

📄 PDF Abstract BibTeX arXiv:2011.13831

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Orthogonal greedy algorithm for linear operator learning with shallow neural network

2025-01-06 · Ye Lin, Jiwei Jia, Young Ju Lee, Ran Zhang

Greedy algorithms, particularly the orthogonal greedy algorithm (OGA), have proven effective in training shallow neural networks for fitting functions and solving partial differential equations (PDEs). In this paper, we …

Operator learning

Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data

2025-10-24 · Hancheng Min, Zhihui Zhu, René Vidal arxiv

Among many mysteries behind the success of deep networks lies the exceptional discriminative power of their learned representations as manifested by the intriguing Neural Collapse (NC) phenomenon, where simple feature st…

Reduced Order Modeling with Shallow Recurrent Decoder Networks

2025-02-15 · Matteo Tomasetto, Jan P. Williams, Francesco Braghin, Andrea Manzoni 외

Reduced Order Modeling is of paramount importance for efficiently inferring high-dimensional spatio-temporal fields in parametric contexts, enabling computationally tractable parametric analyses, uncertainty quantificati…

Computational EfficiencyDecoderDimensionality ReductionUncertainty Quantification

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

2025-07-08 · Yisu Zhang, Chenjie Cao, Chaohui Yu, Jianke Zhu

Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or tempor…

Robustness Reprogramming for Representation Learning

2024-10-06 · Zhichao Hou, MohamadAli Torkamani, Hamid Krim, Xiaorui Liu

This work tackles an intriguing and fundamental open challenge in representation learning: Given a well-trained deep learning model, can it be reprogrammed to enhance its robustness against adversarial or noisy input per…

Deep LearningRepresentation Learning