paper-with-me

Papers

Dual-Flow Reinforcement Learning with State-Aware Exploration

2026-06-29 · Qijun Li, Zheng Fu, Qi Song, Yifei He, Weitao Zhou, Kun Jiang, Diange Yang arxiv

In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging. Existing value estimation methods using unimodal Gaussians restrict expressiveness and yield biased estimates. Recent generative policies can represent multimodal actions but often collapse to a few modes and under-explore high-value areas of the action space. Motivated by these challenges, we propose Dual-Flow RL, a unified actor-critic framework that jointly models a continuous return distribution and a multimodal policy distribution using conditional flow matching (CFM). This design supports reliable value estimation and sustained multimodal exploration. To further enhance exploration, we introduce an Entropy-Covariance Exploration Regulator (ECER) that enables state-aware exploration regulation leveraging policy entropy and action-uncertainty covariance. Experiments on DeepMind Control Suite and Humanoid-Bench show that Dual-Flow RL achieves state-of-the-art performance on most tasks, significantly outperforming prior diffusion-based and flow-based methods.

📄 PDF Abstract BibTeX arXiv:2606.29820

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning

2026-01-31 · Xu Pan, Zhenglin Wan, Xingrui Yu, Xianwei Zheng 외 arxiv

Vision-Language-Action (VLA) models exhibit strong generalization in robotic manipulation, yet reinforcement learning (RL) fine-tuning often degrades robustness under spatial distribution shifts. For flow-matching VLA po…

Representation LearningReinforcement Learning

RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation

2026-02-02 · Entong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek Gupta arxiv

Imitation learning has emerged as an effective approach for bootstrapping sequential decision-making in robotics, achieving strong performance even in high-dimensional dexterous manipulation tasks. Recent behavior clonin…

Reinforcement Learning

SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation

2025-11-26 · Ziyi Chen, Yingnan Guo, Zedong Chu, Minghua Luo 외 arxiv

Embodied navigation that adheres to social norms remains an open research challenge. Our SocialNav is a foundational model for socially-aware navigation with a hierarchical "brain-action" architecture, capable of underst…

Reinforcement Learning

TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

2025-08-06 · Xiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li 외 arxiv

Recent flow matching models for text-to-image generation have achieved remarkable quality, yet their integration with reinforcement learning for human preference alignment remains suboptimal, hindering fine-grained rewar…

Text-to-Image GenerationReinforcement Learning

Multi-Agent Continuous Control with Generative Flow Networks

2024-08-13 · Shuang Luo, Yinchuan Li, Shunyu Liu, Xu Zhang 외

Generative Flow Networks (GFlowNets) aim to generate diverse trajectories from a distribution in which the final states of the trajectories are proportional to the reward, serving as a powerful alternative to reinforceme…

continuous-controlContinuous Control