paper-with-me

홈 › Papers

Expressivity of Quadratic Neural ODEs

2025-04-13 · Joshua Hanson, Maxim Raginsky

This work focuses on deriving quantitative approximation error bounds for neural ordinary differential equations having at most quadratic nonlinearities in the dynamics. The simple dynamics of this model form demonstrates how expressivity can be derived primarily from iteratively composing many basic elementary operations, versus from the complexity of those elementary operations themselves. Like the analog differential analyzer and universal polynomial DAEs, the expressivity is derived instead primarily from the "depth" of the model. These results contribute to our understanding of what depth specifically imparts to the capabilities of deep learning architectures.

📄 PDF Abstract BibTeX arXiv:2504.09385

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Expressivity and Trainability of Quadratic Networks

2021-10-12 · Feng-Lei Fan, Mengzhou Li, Fei Wang, Rongjie Lai 외

Inspired by the diversity of biological neurons, quadratic artificial neurons can play an important role in deep learning models. The type of quadratic neurons of our interest replaces the inner-product operation in the …

Polynormer: Polynomial-Expressive Graph Transformer in Linear Time

2024-03-02 · Chenhui Deng, Zichao Yue, Zhiru Zhang

Graph transformers (GTs) have emerged as a promising architecture that is theoretically more expressive than message-passing graph neural networks (GNNs). However, typical GT models have at least quadratic complexity and…

Node Classification

Quantum CT via Dynamic Interval Encoding and Prior-Balanced QUBO Reconstruction

2026-06-23 · Ao Wang, Yikuang Yuluo, Yujie Liu, Shuangyang Zhong 외 arxiv

Quadratic unconstrained binary optimization (QUBO)-based quantum computed tomography (CT) casts reconstruction as a binary quadratic problem for quantum annealing and hybrid quantum--classical solvers. For grayscale CT, …

SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation

2026-08-17 · Animesh Karnewar, Denis Korzhenkov, Amirhossein Habibian, Mohsen Ghafoorian arxiv

Video Diffusion Transformers (DiTs) spend most of their compute inside the Self-Attention operation, whose cost grows quadratically, $\mathcal{O}(n^2)$, with the number of latent tokens $n$. For the task of video generat…

Video Generation

Erwin: A Tree-based Hierarchical Transformer for Large-scale Physical Systems

2025-02-24 · Maksim Zhdanov, Max Welling, Jan-Willem van de Meent

Large-scale physical systems defined on irregular grids pose significant scalability challenges for deep learning methods, especially in the presence of long-range interactions and multi-scale coupling. Traditional appro…

Computational EfficiencyPDE Surrogate ModelingPhysical Simulations