paper-with-me

홈 › Papers

Re-parameterizing Your Optimizers rather than Architectures

2022-05-30 · Xiaohan Ding, Honghao Chen, Xiangyu Zhang, Kaiqi Huang, Jungong Han, Guiguang Ding

The well-designed structures in neural networks reflect the prior knowledge incorporated into the models. However, though different models have various priors, we are used to training them with model-agnostic optimizers such as SGD. In this paper, we propose to incorporate model-specific prior knowledge into optimizers by modifying the gradients according to a set of model-specific hyper-parameters. Such a methodology is referred to as Gradient Re-parameterization, and the optimizers are named RepOptimizers. For the extreme simplicity of model structure, we focus on a VGG-style plain model and showcase that such a simple model trained with a RepOptimizer, which is referred to as RepOpt-VGG, performs on par with or better than the recent well-designed models. From a practical perspective, RepOpt-VGG is a favorable base model because of its simple structure, high inference speed and training efficiency. Compared to Structural Re-parameterization, which adds priors into models via constructing extra training-time structures, RepOptimizers require no extra forward/backward computations and solve the problem of quantization. We hope to spark further research beyond the realms of model structure design. Code and models \url{https://github.com/DingXiaoH/RepOptimizers}.

📄 PDF Abstract BibTeX arXiv:2205.15242

Code (1)

dingxiaoh/repoptimizers 공식 구현 pytorch

Tasks

Quantization

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling

2026-05-29 · Dmitrii Feoktistov, Timofey Belinsky, Andrey Veprikov, Amir Zainullin 외 arxiv

Sign-based and LMO-inspired optimizers have recently attracted substantial attention in deep learning due to their strong performance and low memory footprint. However, their fixed-magnitude updates can hurt terminal con…

Combining Optimization Methods Using an Adaptive Meta Optimizer

2021-06-19 · Algorithsm MDPI 2021 6 · Nicola Landro, Ignazio Gallo, Riccardo La Grassa

Optimization methods are of great importance for the efficient training of neural networks. There are many articles in the literature that propose particular variants of existing optimizers. In our article, we propose th…

Articles

Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers

2026-02-05 · Artem Riabinin, Andrey Veprikov, Arman Bolatov, Martin Takáč 외 arxiv

We study adaptive learning rate scheduling for norm-constrained optimizers (e.g., Muon and Lion). We introduce a generalized smoothness assumption under which local curvature decreases with the suboptimality gap and empi…

Mixing ADAM and SGD: a Combined Optimization Method

2020-11-16 · Nicola Landro, Ignazio Gallo, Riccardo La Grassa

Optimization methods (optimizers) get special attention for the efficient training of neural networks in the field of deep learning. In literature there are many papers that compare neural models trained with the use of …

Document ClassificationStochastic Optimizationvalid

pCON: Polarimetric Coordinate Networks for Neural Scene Representations

2023-01-01 · CVPR 2023 1 · Henry Peters, Yunhao Ba, Achuta Kadambi

Neural scene representations have achieved great success in parameterizing and reconstructing images, but current state of the art models are not optimized with the preservation of physical quantities in mind. While …