paper-with-me

홈 › Papers

State Relevance for Off-Policy Evaluation

2021-09-13 · Simon P. Shen, Yecheng Jason Ma, Omer Gottesman, Finale Doshi-Velez

Importance sampling-based estimators for off-policy evaluation (OPE) are valued for their simplicity, unbiasedness, and reliance on relatively few assumptions. However, the variance of these estimators is often high, especially when trajectories are of different lengths. In this work, we introduce Omitting-States-Irrelevant-to-Return Importance Sampling (OSIRIS), an estimator which reduces variance by strategically omitting likelihood ratios associated with certain states. We formalize the conditions under which OSIRIS is unbiased and has lower variance than ordinary importance sampling, and we demonstrate these properties empirically.

📄 PDF Abstract BibTeX arXiv:2109.06310

Code (1)

dtak/osiris 공식 구현 pytorch

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Off-policy Evaluation with Deeply-abstracted States

2024-06-27 · Meiling Hao, Pingfan Su, Liyuan Hu, Zoltan Szabo 외

Off-policy evaluation (OPE) is crucial for assessing a target policy's impact offline before its deployment. However, achieving accurate OPE in large state spaces remains challenging. This paper studies state abstraction…

Off-policy evaluation

Multimodal Label Relevance Ranking via Reinforcement Learning

2024-07-18 · Taian Guo, Taolin Zhang, Haoqian Wu, Hanjun Li 외

Conventional multi-label recognition methods often focus on label confidence, frequently overlooking the pivotal role of partial order relations consistent with human preference. To resolve these issues, we introduce a n…

reinforcement-learningReinforcement Learning

Quantifying the Relevance of Youth Research Cited in the US Policy Documents

2025-03-06 · Miftahul Jannat Mokarrama, Hamed Alhoori

In recent years, there has been a growing concern and emphasis on conducting research beyond academic or scientific research communities, benefiting society at large. A well-known approach to measuring the impact of rese…

Articles

Pluralistic Off-policy Evaluation and Alignment

2025-09-15 · Chengkai Huang, Junda Wu, Zhouhang Xie, Yu Xia 외 arxiv

Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that diffe…

Response Generation

SAGE: Scalable AI Governance & Evaluation

2026-02-08 · Benjamin Le, Xueying Lu, Nick Stern, Wenqiong Liu 외 arxiv

Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems. Whil…