paper-with-me

홈 › Papers

Cross-Validated Off-Policy Evaluation

2024-05-24 · Matej Cief, Branislav Kveton, Michal Kompan

We study estimator selection and hyper-parameter tuning in off-policy evaluation. Although cross-validation is the most popular method for model selection in supervised learning, off-policy evaluation relies mostly on theory, which provides only limited guidance to practitioners. We show how to use cross-validation for off-policy evaluation. This challenges a popular belief that cross-validation in off-policy evaluation is not feasible. We evaluate our method empirically and show that it addresses a variety of use cases.

📄 PDF Abstract BibTeX arXiv:2405.15332

Code (1)

navarog/cross-validated-ope 공식 구현 pytorch

Tasks

Model SelectionOff-policy evaluation

Similar Papers 제목 키워드 기반

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

2026-05-29 · Xiaobo Wang, Tong Wu, Min Tang, Jiaqi Li 외 arxiv

Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation or judge models. It is dramatically wors…

Policy-Grounded Safety Evaluation of 20 Large Language Models

2025-07-19 · Juan Manuel Contreras arxiv

As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. This paper introduces Aymara AI, a programmatic platform for generating a…

Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation

2026-06-15 · Adrian Ramlal, Yuhao Chen, John S. Zelek arxiv

Realistic visual simulation of food manipulation requires accurate material parameters, yet these are difficult to measure directly and vary across the heterogeneous regions of a single food item. We address the inverse …

Reinforcement Learning

Policy Learning with a Natural Language Action Space: A Causal Approach

2025-02-24 · Bohan Zhang, Yixin Wang, Paramveer S. Dhillon

This paper introduces a novel causal framework for multi-stage decision-making in natural language action spaces where outcomes are only observed after a sequence of actions. While recent approaches like Proximal Policy …

Decision MakingQ-Learning

POLICY DRIVEN GENERATIVE ADVERSARIAL NETWORKS FOR ACCENTED SPEECH GENERATION

2018-01-01 · ICLR 2018 1 · Prannay Khosla, Preethi Jyothi, Vinay P. Namboodiri, Mukundhan Srinivasan

In this paper, we propose the generation of accented speech using generative adversarial networks. Through this work we make two main contributions a) The ability to condition latent representations while generating real…

Speech Synthesis