paper-with-me

홈 › Papers

Low-Variance and Zero-Variance Baselines for Extensive-Form Games

2019-07-22 · ICML 2020 1 · Trevor Davis, Martin Schmid, Michael Bowling

Extensive-form games (EFGs) are a common model of multi-agent interactions with imperfect information. State-of-the-art algorithms for solving these games typically perform full walks of the game tree that can prove prohibitively slow in large games. Alternatively, sampling-based methods such as Monte Carlo Counterfactual Regret Minimization walk one or more trajectories through the tree, touching only a fraction of the nodes on each iteration, at the expense of requiring more iterations to converge due to the variance of sampled values. In this paper, we extend recent work that uses baseline estimates to reduce this variance. We introduce a framework of baseline-corrected values in EFGs that generalizes the previous work. Within our framework, we propose new baseline functions that result in significantly reduced variance compared to existing techniques. We show that one particular choice of such a function --- predictive baseline --- is provably optimal under certain sampling schemes. This allows for efficient computation of zero-variance value estimates even along sampled trajectories.

📄 PDF Abstract BibTeX arXiv:1907.09633

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualForm

Similar Papers 제목 키워드 기반

Variance Reduction in Monte Carlo Counterfactual Regret Minimization (VR-MCCFR) for Extensive Form Games using Baselines

2018-09-09 · Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik 외

Learning strategies for imperfect information games from samples of interaction is a challenging problem. A common method for this setting, Monte Carlo Counterfactual Regret Minimization (MCCFR), can have slow long-term …

counterfactualFormReinforcement Learning

No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping

2025-09-26 · Thanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Lai 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful framework for improving the reasoning abilities of Large Language Models (LLMs). However, current methods such as GRPO rely only on problems where the m…

Reinforcement Learning

Refining Adaptive Zeroth-Order Optimization at Ease

2025-02-03 · Yao Shu, Qixin Zhang, Kun He, Zhongxiang Dai

Recently, zeroth-order (ZO) optimization plays an essential role in scenarios where gradient information is inaccessible or unaffordable, such as black-box systems and resource-constrained environments. While existing ad…

Adversarial Attack

Zeroth-Order Stochastic Variance Reduction for Nonconvex Optimization

2018-05-25 · NeurIPS 2018 12 · Sijia Liu, Bhavya Kailkhura, Pin-Yu Chen, Pai-Shun Ting 외

As application demands for zeroth-order (gradient-free) optimization accelerate, the need for variance reduced and faster converging approaches is also intensifying. This paper addresses these challenges by presenting: a…

Material ClassificationStochastic Optimization

Embedding-perturbed Exploration Preference Optimization for Flow Models

2026-05-15 · Sujie Hu, Chubin Chen, Jiashu Zhu, Jiahong Wu 외 arxiv

Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-based optimization frameworks (e.g., GRPO) face a critical limitatio…

Reinforcement Learning