paper-with-me

홈 › Papers

Inference on Optimal Dynamic Policies via Softmax Approximation

2023-03-08 · Qizhao Chen, Morgane Austern, Vasilis Syrgkanis

Estimating optimal dynamic policies from offline data is a fundamental problem in dynamic decision making. In the context of causal inference, the problem is known as estimating the optimal dynamic treatment regime. Even though there exists a plethora of methods for estimation, constructing confidence intervals for the value of the optimal regime and structural parameters associated with it is inherently harder, as it involves non-linear and non-differentiable functionals of unknown quantities that need to be estimated. Prior work resorted to sub-sample approaches that can deteriorate the quality of the estimate. We show that a simple soft-max approximation to the optimal treatment regime, for an appropriately fast growing temperature parameter, can achieve valid inference on the truly optimal regime. We illustrate our result for a two-period optimal dynamic regime, though our approach should directly extend to the finite horizon case. Our work combines techniques from semi-parametric inference and $g$-estimation, together with an appropriate triangular array central limit theorem, as well as a novel analysis of the asymptotic influence and asymptotic bias of softmax approximations.

📄 PDF Abstract BibTeX arXiv:2303.04416

Code (1)

syrgkanislab/optimal_dynamic_regime 공식 구현

Tasks

Causal InferenceDecision Makingvalid

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Revisiting the Softmax Bellman Operator: New Benefits and New Perspective

2018-12-02 · Zhao Song, Ronald E. Parr, Lawrence Carin

The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of th…

Atari GamesQ-LearningReinforcement LearningReinforcement Learning (RL)

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map

2024-11-16 · Yuhong Chou, Man Yao, Kexin Wang, Yuqi Pan 외

Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. Howe…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling

Computational Hardness of Reinforcement Learning with Partial $q^π$-Realizability

2025-10-24 · Shayan Karimi, Xiaoqi Tan arxiv

This paper investigates the computational complexity of reinforcement learning in a novel linear function approximation regime, termed partial $q^π$-realizability. In this framework, the objective is to learn an $ε$-opti…

Reinforcement Learning

CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption

2026-02-02 · Hanjun Park, Byeongseo Min, Jiheon Woo, Min-Wook Jeong 외 arxiv

Homomorphic encryption (HE) is a prominent framework for privacy-preserving machine learning, enabling inference directly on encrypted data. However, evaluating softmax, a core component of transformer architectures, rem…

Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods

2023-10-04 · Sara Klein, Simon Weissmann, Leif Döring

Markov Decision Processes (MDPs) are a formal framework for modeling and solving sequential decision-making problems. In finite-time horizons such problems are relevant for instance for optimal stopping or specific suppl…

Decision MakingPolicy Gradient MethodsSequential Decision Making