paper-with-me

홈 › Papers

On the training and generalization of deep operator networks

2023-09-02 · SangHyun Lee, Yeonjong Shin

We present a novel training method for deep operator networks (DeepONets), one of the most popular neural network models for operators. DeepONets are constructed by two sub-networks, namely the branch and trunk networks. Typically, the two sub-networks are trained simultaneously, which amounts to solving a complex optimization problem in a high dimensional space. In addition, the nonconvex and nonlinear nature makes training very challenging. To tackle such a challenge, we propose a two-step training method that trains the trunk network first and then sequentially trains the branch network. The core mechanism is motivated by the divide-and-conquer paradigm and is the decomposition of the entire complex training task into two subtasks with reduced complexity. Therein the Gram-Schmidt orthonormalization process is introduced which significantly improves stability and generalization ability. On the theoretical side, we establish a generalization error estimate in terms of the number of training data, the width of DeepONets, and the number of input and output sensors. Numerical examples are presented to demonstrate the effectiveness of the two-step training method, including Darcy flow in heterogeneous porous media.

📄 PDF Abstract BibTeX arXiv:2309.01020

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Test-time Generalization for Physics through Neural Operator Splitting

2026-01-31 · Louis Serrano, Jiequn Han, Edouard Oyallon, Shirley Ho 외 arxiv

Neural operators have shown promise in learning solution maps of partial differential equations (PDEs), but they often struggle to generalize when test inputs lie outside the training distribution, such as novel initial …

Zero-shot Generalization

Graph In-Context Operator Networks for Generalizable Spatiotemporal Prediction

2026-03-13 · Chenghan Wu, Zongmin Yu, Boai Sun, Liu Yang arxiv

In-context operator learning enables neural networks to infer solution operators from contextual examples without weight updates. While prior work has demonstrated the effectiveness of this paradigm in leveraging vast da…

Generalization Bounds and Statistical Guarantees for Multi-Task and Multiple Operator Learning with MNO Networks

2026-04-02 · Adrien Weihs, Hayden Schaeffer arxiv

Multiple operator learning concerns learning operator families $\{G[α]:U\to V\}_{α\in W}$ indexed by an operator descriptor $α$. Training data are collected hierarchically by sampling operator instances $α$, then input f…

Multi-Operator Few-Shot Learning for Generalization Across PDE Families

2025-08-02 · Yile Li, Shandian Zhe arxiv

Learning solution operators for partial differential equations (PDEs) has become a foundational task in scientific machine learning. However, existing neural operator methods require abundant training data for each speci…

Few-Shot Learning

Generalization Guarantees for Multi-Input Neural Operator Learning in Sobolev Spaces

2026-06-16 · Yahong Yang, Zecheng Zhang, Wei Zhu, Wenjing Liao 외 arxiv

We develop approximation and generalization error estimates for multi-input neural operators, with the output error measured in Sobolev norms. In contrast to standard operator-learning settings with a single input functi…