paper-with-me

홈 › Papers

Adaptive Learning via Off-Model Training and Importance Sampling for Fully Non-Markovian Optimal Stochastic Control. Complete version

2026-04-14 · Dorival Leão, Alberto Ohashi, Simone Scotti, Adolfo M. D da Silva arxiv

This paper studies continuous-time stochastic control problems whose controlled states are fully non-Markovian and depend on unknown model parameters. Such problems arise naturally in path-dependent stochastic differential equations, rough-volatility hedging, and systems driven by fractional Brownian motion. Building on the discrete skeleton approach developed in earlier work, we propose a Monte Carlo learning methodology for the associated embedded backward dynamic programming equation. Our main contribution is twofold. First, we construct explicit dominating training laws and Radon--Nikodym weights for several representative classes of non-Markovian controlled systems. This yields an off-model training architecture in which a fixed synthetic dataset is generated under a reference law, while the dynamic programming operators associated with a target model are recovered by importance sampling. Second, we use this structure to design an adaptive update mechanism under parametric model uncertainty, so that repeated recalibration can be performed by reweighting the same training sample rather than regenerating new trajectories. For fixed parameters, we establish non-asymptotic error bounds for the approximation of the embedded dynamic programming equation via deep neural networks. For adaptive learning, we derive quantitative estimates that separate Monte Carlo approximation error from model-risk error. Numerical experiments illustrate both the off-model training mechanism and the adaptive importance-sampling update in structured linear-quadratic examples.

📄 PDF Abstract BibTeX arXiv:2604.13147

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Importance Sampling Policy Evaluation with an Estimated Behavior Policy

2018-06-04 · Josiah P. Hanna, Scott Niekum, Peter Stone

We consider the problem of off-policy evaluation in Markov decision processes. Off-policy evaluation is the task of evaluating the expected return of one policy with data generated by a different, behavior policy. Import…

Off-policy evaluation

Online covariance estimation for stochastic gradient descent under Markovian sampling

2023-08-03 · Abhishek Roy, Krishnakumar Balasubramanian

We investigate the online overlapping batch-means covariance estimator for Stochastic Gradient Descent (SGD) under Markovian sampling. Convergence rates of order $O\big(\sqrt{d}\,n^{-1/8}(\log n)^{1/4}\big)$ and $O\big(\…

regression

Non-Asymptotic Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes

2025-04-25 · Zaiwei Chen

This work presents the first finite-time analysis for the last-iterate convergence of average-reward Q-learning with an asynchronous implementation. A key feature of the algorithm we study is the use of adaptive stepsize…

Q-Learning

Importance sampling for partially observed temporal epidemic models

2018-08-15

We present an importance sampling algorithm that can produce realisations of Markovian epidemic models that exactly match observations, taken to be the number of a single event type over a period of time. The importance …

Online learning of quantum processes

2024-06-06 · Asad Raza, Matthias C. Caro, Jens Eisert, Sumeet Khatri

Among recent insights into learning quantum states, online learning and shadow tomography procedures are notable for their ability to accurately predict expectation values even of adaptively chosen observables. In contra…