paper-with-me

홈 › Papers

Accuracy of Discretely Sampled Stochastic Policies in Continuous-time Reinforcement Learning

2025-03-13 · Yanwei Jia, Du Ouyang, Yufei Zhang

Stochastic policies are widely used in continuous-time reinforcement learning algorithms. However, executing a stochastic policy and evaluating its performance in a continuous-time environment remain open challenges. This work introduces and rigorously analyzes a policy execution framework that samples actions from a stochastic policy at discrete time points and implements them as piecewise constant controls. We prove that as the sampling mesh size tends to zero, the controlled state process converges weakly to the dynamics with coefficients aggregated according to the stochastic policy. We explicitly quantify the convergence rate based on the regularity of the coefficients and establish an optimal first-order convergence rate for sufficiently regular coefficients. Additionally, we show that the same convergence rates hold with high probability concerning the sampling noise, and further establish a $1/2$-order almost sure convergence when the volatility is not controlled. Building on these results, we analyze the bias and variance of various policy evaluation and policy gradient estimators based on discrete-time observations. Our results provide theoretical justification for the exploratory stochastic control framework in [H. Wang, T. Zariphopoulou, and X.Y. Zhou, J. Mach. Learn. Res., 21 (2020), pp. 1-34].

📄 PDF Abstract BibTeX arXiv:2503.09981

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pricing variance swaps with stochastic volatility and stochastic interest rate under full correlation structure

2016-10-30 · Teh Raihana Nazirah Roslan, Wenjun Zhang, Jiling Cao

This paper considers the case of pricing discretely-sampled variance swaps under the class of equity-interest rate hybridization. Our modeling framework consists of the equity which follows the dynamics of the Heston sto…

Pricing variance swaps in a hybrid model of stochastic volatility and interest rate with regime-switching

2016-03-28

In this paper, we consider the problem of pricing discretely-sampled variance swaps based on a hybrid model of stochastic volatility and stochastic interest rate with regime-switching. Our modelling framework extends the…

Naive Markowitz Policies

2022-12-14 · Lin Chen, Xun Yu Zhou

We study a continuous-time Markowitz mean-variance portfolio selection model in which a naive agent, unaware of the underlying time-inconsistency, continuously reoptimizes over time. We define the resulting naive policie…

Discretization and Machine Learning Approximation of BSDEs with a Constraint on the Gains-Process

2020-02-07 · Idris Kharroubi, Thomas Lim, Xavier Warin

We study the approximation of backward stochastic differential equations (BSDEs for short) with a constraint on the gains process. We first discretize the constraint by applying a so-called facelift operator at times of …

BIG-bench Machine Learning

Pricing timer options and variance derivatives with closed-form partial transform under the 3/2 model

2015-04-30

Most of the empirical studies on stochastic volatility dynamics favor the 3/2 specification over the square-root (CIR) process in the Heston model. In the context of option pricing, the 3/2 stochastic volatility model is…

Form