paper-with-me

홈 › Papers

Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity

2020-11-05 · Tanmay Gangwani, Jian Peng, Yuan Zhou

Quality-Diversity (QD) is a concept from Neuroevolution with some intriguing applications to Reinforcement Learning. It facilitates learning a population of agents where each member is optimized to simultaneously accumulate high task-returns and exhibit behavioral diversity compared to other members. In this paper, we build on a recent kernel-based method for training a QD policy ensemble with Stein variational gradient descent. With kernels based on $f$-divergence between the stationary distributions of policies, we convert the problem to that of efficient estimation of the ratio of these stationary distributions. We then study various distribution ratio estimators used previously for off-policy evaluation and imitation and re-purpose them to compute the gradients for policies in an ensemble such that the resultant population is diverse and of high-quality.

📄 PDF Abstract BibTeX arXiv:2011.02614

Code (1)

tgangwani/QDAgents 공식 구현 pytorch

Tasks

DiversityOff-policy evaluation

Similar Papers 제목 키워드 기반

Identification and Estimation of a Semiparametric Logit Model using Network Data

2023-10-11 · Brice Romuald Gueyap Kounga

This paper studies the identification and estimation of a semiparametric binary network model in which the unobserved social characteristic is endogenous, that is, the unobserved individual characteristic influences both…

A Theory of the Distortion-Perception Tradeoff in Wasserstein Space

2021-07-06 · NeurIPS 2021 12 · Dror Freirich, Tomer Michaeli, Ron Meir

The lower the distortion of an estimator, the more the distribution of its outputs generally deviates from the distribution of the signals it attempts to estimate. This phenomenon, known as the perception-distortion trad…

FormImage RestorationOpen-Ended Question Answering

Bounding Wasserstein distance with couplings

2021-12-06 · pproximateinference AABI Symposium 2022 2 · Niloy Biswas, Lester Mackey

Markov chain Monte Carlo (MCMC) provides asymptotically consistent estimates of intractable posterior expectations as the number of iterations tends to infinity. However, in large data applications, MCMC can be computati…

regression

Harnessing Agent Skills: Architectural Patterns and a Reference Architecture for Skill-Mediated LLM Agents

2026-05-29 · Boming Xia, Liming Zhu, Zhenchang Xing, Qinghua Lu 외 arxiv

Agent skills externalise reusable agent-facing behavioural knowledge and guidance as persistent artefacts that can be discovered, activated, and interpreted by LLM agents. Although a skill artefact is static at rest, its…

Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents

2024-05-21 · San Kim, Gary Geunbae Lee

Recent advancements in open-domain dialogue systems have been propelled by the emergence of high-quality large language models (LLMs) and various effective training methodologies. Nevertheless, the presence of toxicity w…