paper-with-me

Papers

Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models

2025-12-30 · Lars van der Laan, Aurelien Bibaut, Nathan Kallus arxiv

In many sequential decision-making problems, researchers observe actions but not the rewards that drive behavior, yet still wish to evaluate and compare counterfactual policies. Inverse reinforcement learning (IRL) and dynamic discrete choice (DDC) models address this setting by positing an optimality model that links latent rewards to observed actions. Existing flexible IRL methods allow rich reward representations but typically do not provide valid inference, whereas classical DDC methods support inference only under restrictive parametric structure. We develop a semiparametric framework for debiased inverse reinforcement learning in maximum-entropy IRL and Gumbel-shock DDC models. Our key identification result is that the log-behavior policy can be treated as a pseudo-reward: it point-identifies policy value differences and, under a normalization constraint, the reward itself. This reduces inference on reward-dependent estimands to inference on smooth functionals of the behavior policy and transition kernel. We establish pathwise differentiability, derive efficient influence functions, and construct automatic debiased machine-learning estimators that permit flexible nuisance estimation while attaining $\sqrt{n}$-consistency, asymptotic normality, and semiparametric efficiency. The result is a computationally tractable framework for valid uncertainty quantification in flexible IRL and DDC models.

📄 PDF Abstract BibTeX arXiv:2512.24407

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models

2026-05-29 · Enoch Hyunwook Kang arxiv

In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given offline data generated by an expert, can…

Reinforcement LearningOffline RL

A tutorial on recursive models for analyzing and predicting path choice behavior

2019-05-02 · Maëlle Zimmermann, Emma Frejinger

The problem at the heart of this tutorial consists in modeling the path choice behavior of network users. This problem has been extensively studied in transportation science, where it is known as the route choice problem…

Discrete Choice ModelsReinforcement Learning

A Data-Driven State Aggregation Approach for Dynamic Discrete Choice Models

2023-04-11 · Sinong Geng, Houssam Nassif, Carlos A. Manzanares

We study dynamic discrete choice models, where a commonly studied problem involves estimating parameters of agent reward functions (also known as "structural" parameters), using agent behavioral data. Maximum likelihood …

Discrete Choice Models

An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model

2025-02-19 · Enoch H. Kang, Hema Yoganarasimhan, Lalit Jain

We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in machine learning. The objective is to reco…

A deep inverse reinforcement learning approach to route choice modeling with context-dependent rewards

2022-06-18 · Zhan Zhao, Yuebing Liang

Route choice modeling is a fundamental task in transportation planning and demand forecasting. Classical methods generally adopt the discrete choice model (DCM) framework with linear utility functions and high-level rout…

Computational EfficiencyDemand ForecastingImitation Learningreinforcement-learning+1