paper-with-me

홈 › Papers

A First-Order Mean Field Control Analysis of Transformer Layers under Cross-Entropy Training

2026-06-22 · Cheng Huan, Hongwei Yuan arxiv

We study Transformer-type residual layers under cross-entropy training through a continuous-depth mean field control viewpoint. Depth is treated as time, layer parameters as controls, and the residual Transformer recursion as an explicit Euler scheme for a controlled hidden-state flow. For fixed controls, we prove an $O(\varepsilon)$ pathwise approximation of finite-depth trajectories by the continuous flow and combine this with high-probability sampling bounds for the empirical cross-entropy risk. We formulate the limiting population problem as a first-order transport control problem for the law of hidden states and derive a Pontryagin condition whose terminal adjoint contains the softmax residual. We also give finite-class and metric-entropy uniform estimates, compare optimal values, and discuss existence, stability, continuous-to-discrete recovery, initialization, and range estimates for continuous minimizers.

📄 PDF Abstract BibTeX arXiv:2606.23235

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Convergence Analysis of Machine Learning Algorithms for the Numerical Solution of Mean Field Control and Games: II -- The Finite Horizon Case

2019-08-05 · René Carmona, Mathieu Laurière

We propose two numerical methods for the optimal control of McKean-Vlasov dynamics in finite time horizon. Both methods are based on the introduction of a suitable loss function defined over the parameters of a neural ne…

Backstepping Mean-Field Density Control for Large-Scale Heterogeneous Nonlinear Stochastic Systems

2021-09-01 · Tongjia Zheng, Qing Han, Hai Lin

This work studies the problem of controlling the mean-field density of large-scale stochastic systems, which has applications in various fields such as swarm robotics. Recently, there is a growing amount of literature th…

Computational Efficiency

Scalable Task-Driven Robotic Swarm Control via Collision Avoidance and Learning Mean-Field Control

2022-09-15 · Kai Cui, Mengguang Li, Christian Fabian, Heinz Koeppl

In recent years, reinforcement learning and its multi-agent analogue have achieved great success in solving various complex control problems. However, multi-agent reinforcement learning remains challenging both in its th…

Collision AvoidanceMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+1

Hodge Decomposition of the wall shear stress vector fields characterizing biological flows

2018-11-21

A discrete boundary-sensitive Hodge decomposition is proposed as a central tool for the analysis of wall shear stress (WSS) vector fields in aortic blood flows. The method is based on novel results for the smooth and dis…

Score-based Neural Ordinary Differential Equations for Computing Mean Field Control Problems

2024-09-24 · Mo Zhou, Stanley Osher, Wuchen Li

Classical neural ordinary differential equations (ODEs) are powerful tools for approximating the log-density functions in high-dimensional spaces along trajectories, where neural networks parameterize the velocity fields…