paper-with-me

홈 › Papers

Scheduled Policy Optimization for Natural Language Communication with Intelligent Agents

2018-06-16 · Wenhan Xiong, Xiaoxiao Guo, Mo Yu, Shiyu Chang, Bo-Wen Zhou, William Yang Wang

We investigate the task of learning to follow natural language instructions by jointly reasoning with visual observations and language inputs. In contrast to existing methods which start with learning from demonstrations (LfD) and then use reinforcement learning (RL) to fine-tune the model parameters, we propose a novel policy optimization algorithm which dynamically schedules demonstration learning and RL. The proposed training paradigm provides efficient exploration and better generalization beyond existing methods. Comparing to existing ensemble models, the best single model based on our proposed method tremendously decreases the execution error by over 50% on a block-world environment. To further illustrate the exploration strategy of our RL algorithm, We also include systematic studies on the evolution of policy entropy during training.

📄 PDF Abstract BibTeX arXiv:1806.06187

Code (3)

xwhan/walk_the_blocks 공식 구현 pytorch
clic-lab/ciff pytorch
lil-lab/ciff pytorch

Tasks

Efficient Explorationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Dynamic Scheduling for Federated Edge Learning with Streaming Data

2023-05-02 · Chung-Hsuan Hu, Zheng Chen, Erik G. Larsson

In this work, we consider a Federated Edge Learning (FEEL) system where training data are randomly generated over time at a set of distributed edge devices with long-term energy constraints. Due to limited communication …

Scheduling

Federated Reinforcement Learning with Constraint Heterogeneity

2024-05-06 · Hao Jin, Liangyu Zhang, Zhihua Zhang

We study a Federated Reinforcement Learning (FedRL) problem with constraint heterogeneity. In our setting, we aim to solve a reinforcement learning problem with multiple constraints while $N$ training agents are located …

Language ModelingLanguage ModellingLarge Language ModelPolicy Gradient Methods+2

Not only where, But when: Temporal Scheduling for RLVR

2026-05-25 · Jinghao Zhang, Ruilin Li, Feng Zhao, Jiaqi Wang arxiv

Reinforcement learning with verifiable rewards (RLVR) has become a core technique for post-training of Large Language Models (LLMs). While policy optimization is driven by all sampled tokens under a globally broadcast sc…

Reinforcement Learning

Energy-Aware Analog Aggregation for Federated Learning with Redundant Data

2019-11-01 · Yuxuan Sun, Sheng Zhou, Deniz Gündüz

Federated learning (FL) enables workers to learn a model collaboratively by using their local data, with the help of a parameter server (PS) for global model aggregation. The high communication cost for periodic model up…

Federated LearningScheduling

Personalized Execution Time Optimization for the Scheduled Jobs

2022-03-11 · Yang Liu, Juan Wang, Zhengxing Chen, Ian Fox 외

Scheduled batch jobs have been widely used on the asynchronous computing platforms to execute various enterprise applications, including the scheduled notifications and the candidate pre-computation for the modern recomm…

Learning-To-RankRecommendation SystemsScheduling