paper-with-me

홈 › Papers

Generalized Distributional Alignment Games for Unbiased Answer-Level Fine-Tuning

2026-05-04 · Mehryar Mohri, Jon Schneider, Yutao Zhong arxiv

The Distributional Alignment Game framework provides a powerful variational perspective on Answer-Level Fine-Tuning (ALFT). However, standard algorithms for these games rely on estimating logarithmic rewards from small batches, introducing a systematic bias due to Jensen's inequality that can destabilize training. In this paper, we systematically resolve this structural estimation bias. First, we generalize the alignment game to arbitrary Bregman divergences, showing that for a family of geometries inducing polynomial rewards, we can construct provably exact and unbiased estimators using U-statistics. Second, for the canonical KL divergence game where an exact solution is impossible, we derive a globally robust minimax polynomial estimator that is provably optimal, achieving the fundamental statistical error limit of $Θ(1/K^2)$, which we establish via the Ditzian-Totik theorem. Finally, we synthesize these two approaches to propose a novel Variance-Optimal Augmented Polynomial Optimization Program (AQP) Estimator, proving that by systematically reducing variance, our method achieves not only optimal bias but also provably accelerated game convergence, leading to more efficient and stable training with zero online computational overhead.

📄 PDF Abstract BibTeX arXiv:2605.02435

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Distributional Alignment Games for Answer-Level Fine-Tuning

2026-04-29 · Mehryar Mohri, Jon Schneider, Yifan Wu arxiv

We focus on the problem of \emph{Answer-Level Fine-Tuning} (ALFT), where the goal is to optimize a language model based on the correctness or properties of its final answers, rather than the specific reasoning traces use…

Mathematical Reasoning

Adjusted Wasserstein Distributionally Robust Estimator in Statistical Learning

2023-03-27 · Yiling Xie, Xiaoming Huo

We propose an adjusted Wasserstein distributionally robust estimator -- based on a nonlinear transformation of the Wasserstein distributionally robust (WDRO) estimator in statistical learning. The classic WDRO estimator …

regression

A Robust Quantile Huber Loss With Interpretable Parameter Adjustment In Distributional Reinforcement Learning

2024-01-04 · Parvin Malekzadeh, Konstantinos N. Plataniotis, Zissis Poulos, Zeyu Wang

Distributional Reinforcement Learning (RL) estimates return distribution mainly by learning quantile values via minimizing the quantile Huber loss function, entailing a threshold parameter often selected heuristically or…

Atari GamesDistributional Reinforcement LearningReinforcement Learning (RL)

Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models

2026-02-02 · Yingsha Xie, Tiansheng Huang, Enneng Yang, Rui Min 외 arxiv

Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually constructed by distilling safety reasoning tr…

Computing Nash Equilibria in Generalized Interdependent Security Games

2014-12-01 · NeurIPS 2014 12 · Hau Chan, Luis E. Ortiz

We study the computational complexity of computing Nash equilibria in generalized interdependent-security (IDS) games. Like traditional IDS games, originally introduced by economists and risk-assessment experts Heal and …