paper-with-me

Papers

$α$-GAN: Convergence and Estimation Guarantees

2022-05-12 · Gowtham R. Kurri, Monica Welfert, Tyler Sypherd, Lalitha Sankar

We prove a two-way correspondence between the min-max optimization of general CPE loss function GANs and the minimization of associated $f$-divergences. We then focus on $\alpha$-GAN, defined via the $\alpha$-loss, which interpolates several GANs (Hellinger, vanilla, Total Variation) and corresponds to the minimization of the Arimoto divergence. We show that the Arimoto divergences induced by $\alpha$-GAN equivalently converge, for all $\alpha\in \mathbb{R}_{>0}\cup\{\infty\}$. However, under restricted learning models and finite samples, we provide estimation bounds which indicate diverse GAN behavior as a function of $\alpha$. Finally, we present empirical results on a toy dataset that highlight the practical utility of tuning the $\alpha$ hyperparameter.

📄 PDF Abstract BibTeX arXiv:2205.06393

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CPE CPE is an effective collaborative metric learning to effectively address the problem of sparse and insufficient preference supervision from the margin distribution point-of-view.

Similar Papers 제목 키워드 기반

Algorithms for ridge estimation with convergence guarantees

2021-04-26 · Wanli Qiao, Wolfgang Polonik

The extraction of filamentary structure from a point cloud is discussed. The filaments are modeled as ridge lines or higher dimensional ridges of an underlying density. We propose two novel algorithms, and provide theore…

Distributed Statistical Estimation and Rates of Convergence in Normal Approximation

2017-04-09 · Stanislav Minsker, Nate Strawn

This paper presents a class of new algorithms for distributed statistical estimation that exploit divide-and-conquer approach. We show that one of the key benefits of the divide-and-conquer strategy is robustness, an imp…

SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees

2026-02-06 · Tianyi Hu, Qingxu Fu, Yanxi Chen, Zhaoyang Liu 외 arxiv

Reinforcement learning (RL) has emerged as the predominant paradigm for training large language model (LLM)-based AI agents. However, existing backbone RL algorithms lack verified convergence guarantees in agentic scenar…

Reinforcement Learning

Statistical, Robustness, and Computational Guarantees for Sliced Wasserstein Distances

2022-10-17 · Sloan Nietert, Ritwik Sadhu, Ziv Goldfeld, Kengo Kato

Sliced Wasserstein distances preserve properties of classic Wasserstein distances while being more scalable for computation and estimation in high dimensions. The goal of this work is to quantify this scalability from th…

Numerical Integration

Fast Rates for the Regret of Offline Reinforcement Learning

2021-01-31 · Yichun Hu, Nathan Kallus, Masatoshi Uehara

We study the regret of reinforcement learning from offline data generated by a fixed behavior policy in an infinite-horizon discounted Markov decision process (MDP). While existing analyses of common approaches, such as …

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)