paper-with-me

홈 › Papers

Layer-wise Derivative Controlled Networks

2026-05-14 · Rowan Martnishn, Sean Anderson arxiv

As machine learning models grow in complexity, they increasingly struggle with three conflicting demands: the need for high accuracy, the requirement for hardware efficiency, and the necessity of functional stability. Traditional architectures often achieve performance at the expense of spiky or unpredictable behavior, where small changes in input lead to massive swings in output -- a critical flaw for real-world deployment in sensitive environments. This paper introduces ChainzRule (CR), a novel neural architecture designed to harmonize these competing goals. ChainzRule replaces standard piecewise-linear activations with a Polynomial Engine governed by Differential Regularization (DREG). Unlike traditional methods that impose global, coarse-grained constraints on a model's Lipschitz constant, DREG acts as a targeted regularization on intermediate derivatives. This approach suppresses extreme sensitivity without attenuating the representational power inherent in the Polynomial Engine. In head-to-head "Fair Fight" benchmarks, ChainzRule outperformed standard models while using 15.5x fewer parameters. On the MNIST dataset, it reduced peak gradient volatility by an average of 23.1%, ensuring a smoother and more predictable manifold. On Yelp Full ordinal regression under explicit DREG regularization, ChainzRule achieves 70.17% accuracy, validating that derivative-aware regularization is compatible with competitive performance on realistic tasks. By embedding gradient awareness into the architecture via DREG, ChainzRule demonstrates that stability and accuracy need not be competing objectives.

📄 PDF Abstract BibTeX arXiv:2605.15463

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Layer-wise Derivative Controlled Networks Achieve Competitive Accuracy and Gradient Stability Across Data Regimes

2026-06-06 · Rowan Martnishn arxiv

Derivative-controlled networks based on ChainzRule (CR) combine cubic polynomial layers with a lightweight forward-mode per-layer Jacobian penalty (DREG). In this second paper of a multi-part series, we evaluate the gene…

Learning to Approximate Adaptive Kernel Convolution on Graphs

2024-01-22 · Jaeyoon Sim, Sooyeon Jeon, InJun Choi, Guorong Wu 외

Various Graph Neural Networks (GNNs) have been successful in analyzing data in non-Euclidean spaces, however, they have limitations such as oversmoothing, i.e., information becomes excessively averaged as the number of h…

Higher-Order Transformer Derivative Estimates for Explicit Pathwise Learning Guarantees

2024-05-26 · Yannick Limmer, Anastasis Kratsios, Xuwei Yang, Raeid Saqur 외

An inherent challenge in computing fully-explicit generalization bounds for transformers involves obtaining covering number estimates for the given transformer class $T$. Crude estimates rely on a uniform upper bound on …

Generalization Bounds

Lifted Proximal Operator Machines

2018-11-05 · Jia Li, Cong Fang, Zhouchen Lin

We propose a new optimization method for training feed-forward neural networks. By rewriting the activation function as an equivalent proximal operator, we approximate a feed-forward neural network by adding the proximal…

Adapting Newton's Method to Neural Networks through a Summary of Higher-Order Derivatives

2023-12-06 · Pierre Wolinski

When training large models, such as neural networks, the full derivatives of order 2 and beyond are usually inaccessible, due to their computational cost. This is why, among the second-order optimization methods, it is v…

Second-order methods