Cross-Validated Off-Policy Evaluation
We study estimator selection and hyper-parameter tuning in off-policy evaluation. Although cross-validation is the most popular method for model selection in supervised learning, off-policy evaluation relies mostly on theory, which provides only limited guidance to practitioners. We show how to use cross-validation for off-policy evaluation. This challenges a popular belief that cross-validation in off-policy evaluation is not feasible. We evaluate our method empirically and show that it addresses a variety of use cases.
Code (1)
Tasks
Model SelectionOff-policy evaluationSimilar Papers 제목 키워드 기반
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation or judge models. It is dramatically wors…
Policy-Grounded Safety Evaluation of 20 Large Language Models
As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. This paper introduces Aymara AI, a programmatic platform for generating a…
Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation
Realistic visual simulation of food manipulation requires accurate material parameters, yet these are difficult to measure directly and vary across the heterogeneous regions of a single food item. We address the inverse …
Reinforcement LearningPolicy Learning with a Natural Language Action Space: A Causal Approach
This paper introduces a novel causal framework for multi-stage decision-making in natural language action spaces where outcomes are only observed after a sequence of actions. While recent approaches like Proximal Policy …
Decision MakingQ-LearningPOLICY DRIVEN GENERATIVE ADVERSARIAL NETWORKS FOR ACCENTED SPEECH GENERATION
In this paper, we propose the generation of accented speech using generative adversarial networks. Through this work we make two main contributions a) The ability to condition latent representations while generating real…
Speech Synthesis