paper-with-me

홈 › Papers

Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints

2025-01-08 · Pavel Kolev, Marin Vlastelica, Georg Martius

While many algorithms for diversity maximization under imitation constraints are online in nature, many applications require offline algorithms without environment interactions. Tackling this problem in the offline setting, however, presents significant challenges that require non-trivial, multi-stage optimization processes with non-stationary rewards. In this work, we present a novel offline algorithm that enhances diversity using an objective based on Van der Waals (VdW) force and successor features, and eliminates the need to learn a previously used skill discriminator. Moreover, by conditioning the value function and policy on a pre-trained Functional Reward Encoding (FRE), our method allows for better handling of non-stationary rewards and provides zero-shot recall of all skills encountered during training, significantly expanding the set of skills learned in prior work. Consequently, our algorithm benefits from receiving a consistently strong diversity signal (VdW), and enjoys more stable and efficient training. We demonstrate the effectiveness of our method in generating diverse skills for two robotic tasks in simulation: locomotion of a quadruped and local navigation with obstacle traversal.

📄 PDF Abstract BibTeX arXiv:2501.04426

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Offline Diversity Maximization Under Imitation Constraints

2023-07-21 · Marin Vlastelica, Jin Cheng, Georg Martius, Pavel Kolev

There has been significant recent progress in the area of unsupervised skill discovery, utilizing various information-theoretic objectives as measures of diversity. Despite these advances, challenges remain: current meth…

D4RLDiversityImitation Learning

Multi-Fidelity Hybrid Reinforcement Learning via Information Gain Maximization

2025-09-18 · Houssem Sifaou, Osvaldo Simeone arxiv

Optimizing a reinforcement learning (RL) policy typically requires extensive interactions with a high-fidelity simulator of the environment, which are often costly or impractical. Offline RL addresses this problem by all…

Reinforcement LearningOffline RL

Fairness Maximization among Offline Agents in Online-Matching Markets

2021-09-18 · Will Ma, Pan Xu, Yifan Xu

Matching markets involve heterogeneous agents (typically from two parties) who are paired for mutual benefit. During the last decade, matching markets have emerged and grown rapidly through the medium of the Internet. Th…

Decision MakingFairness

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

2026-02-11 · Haihui Pan, Yuzhong Hong, Kaichen Zhang, Shaoke Lv 외 arxiv

In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, existing methods often face a fundamental trade-off between these objectives:…

Looking Backward: Retrospective Backward Synthesis for Goal-Conditioned GFlowNets

2024-06-03 · Haoran He, Can Chang, Huazhe Xu, Ling Pan

Generative Flow Networks (GFlowNets), a new family of probabilistic samplers, have demonstrated remarkable capabilities to generate diverse sets of high-reward candidates, in contrast to standard return maximization appr…