paper-with-me

홈 › Papers

Learning controllable dynamics through informative exploration

2025-07-09 · Peter N. Loxley, Friedrich T. Sommer arxiv

Environments with controllable dynamics are usually understood in terms of explicit models. However, such models are not always available, but may sometimes be learned by exploring an environment. In this work, we investigate using an information measure called "predicted information gain" to determine the most informative regions of an environment to explore next. Applying methods from reinforcement learning allows good suboptimal exploring policies to be found, and leads to reliable estimates of the underlying controllable dynamics. This approach is demonstrated by comparing with several myopic exploration approaches.

📄 PDF Abstract BibTeX arXiv:2507.06582

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

An Optimal Policy for Learning Controllable Dynamics by Exploration

2025-12-23 · Peter N. Loxley arxiv

Controllable Markov chains describe the dynamics of sequential decision making tasks and are the central component in optimal control and reinforcement learning. In this work, we give the general form of an optimal polic…

Reinforcement LearningDecision Making

Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning

2026-02-22 · Zhuoxu Huang, Mengxi Jia, Hao Sun, Xuelong Li 외 arxiv

Reinforcement Learning with verifiable rewards (RLVR) has emerged as a primary learning paradigm for enhancing the reasoning capabilities of multi-modal large language models (MLLMs). However, during RL training, the eno…

Reinforcement Learning

Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off

2026-01-19 · Zhaochun Li, Chen Wang, Jionghao Bai, Shisheng Cui 외 arxiv

The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO), training tends to be exploitation driven…

Reinforcement Learning

Towards Empowerment Gain through Causal Structure Learning in Model-Based RL

2025-02-14 · Hongye Cao, Fan Feng, Meng Fang, Shaokang Dong 외

In Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision. Empowerment as an intri…

Causal DiscoveryModel-based Reinforcement Learning

Modeling epidemics on adaptively evolving networks: a data-mining perspective

2015-06-25

The exploration of epidemic dynamics on dynamically evolving ("adaptive") networks poses nontrivial challenges to the modeler, such as the determination of a small number of informative statistics of the detailed network…