paper-with-me

홈 › Papers

Avoid What You Know: Divergent Trajectory Balance for GFlowNets

2026-02-19 · Pedro Dall'Antonia, Tiago da Silva, Daniel Csillag, Salem Lahlou, Diego Mesquita arxiv

Generative Flow Networks (GFlowNets) are a flexible family of amortized samplers trained to generate discrete and compositional objects with probability proportional to a reward function. However, learning efficiency is constrained by the model's ability to rapidly explore diverse high-probability regions during training. To mitigate this issue, recent works have focused on incentivizing the exploration of unvisited and valuable states via curiosity-driven search and self-supervised random network distillation, which tend to waste samples on already well-approximated regions of the state space. In this context, we propose Adaptive Complementary Exploration (ACE), a principled algorithm for the effective exploration of novel and high-probability regions when learning GFlowNets. To achieve this, ACE introduces an exploration GFlowNet explicitly trained to search for high-reward states in regions underexplored by the canonical GFlowNet, which learns to sample from the target distribution. Through extensive experiments, we show that ACE significantly improves upon prior work in terms of approximation accuracy to the target distribution and discovery rate of diverse high-reward states.

📄 PDF Abstract BibTeX arXiv:2602.17827

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance

2026-05-29 · Yuxuan Jiang, Francis Ferraro arxiv

On-Policy Distillation (OPD) improves large language model reasoning by training a student model on trajectories sampled from its own policy under teacher supervision. Although OPD operates on trajectories, its learning …

Monocular Obstacle Avoidance Based on Inverse PPO for Fixed-wing UAVs

2024-11-27 · Haochen Chai, Meimei Su, Yang Lyu, ZhunGa Liu 외

Fixed-wing Unmanned Aerial Vehicles (UAVs) are one of the most commonly used platforms for the burgeoning Low-altitude Economy (LAE) and Urban Air Mobility (UAM), due to their long endurance and high-speed capabilities. …

Collision AvoidanceDeep Reinforcement LearningEdge-computing

Domain-Adaptive Model Merging Across Disconnected Modes

2026-03-06 · Junming Liu, Yusen Zhang, Rongchao Zhang, Wenkai Zhu 외 arxiv

Learning across domains is challenging when data cannot be centralized due to privacy or heterogeneity, which limits the ability to train a single comprehensive model. Model merging provides an appealing alternative by c…

MAXS: Meta-Adaptive Exploration with LLM Agents

2026-01-14 · Jian Zhang, Zhiyuan Wang, Zhangqi Wang, Yu He 외 arxiv

Large Language Model (LLM) Agents exhibit inherent reasoning abilities through the collaboration of multiple tools. However, during agent inference, existing methods often suffer from (i) locally myopic generation, due t…

Computational Efficiency

Towards Better Entity Linking with Multi-View Enhanced Distillation

2023-05-27 · Yi Liu, Yuan Tian, Jianxun Lian, Xinlong Wang 외

Dense retrieval is widely used for entity linking to retrieve entities from large-scale knowledge bases. Mainstream techniques are based on a dual-encoder framework, which encodes mentions and entities independently and …

Entity LinkingKnowledge DistillationRetrieval