paper-with-me

홈 › Papers

DEFT: Diverse Ensembles for Fast Transfer in Reinforcement Learning

2022-09-26 · Simeon Adebola, Satvik Sharma, Kaushik Shivakumar

Deep ensembles have been shown to extend the positive effect seen in typical ensemble learning to neural networks and to reinforcement learning (RL). However, there is still much to be done to improve the efficiency of such ensemble models. In this work, we present Diverse Ensembles for Fast Transfer in RL (DEFT), a new ensemble-based method for reinforcement learning in highly multimodal environments and improved transfer to unseen environments. The algorithm is broken down into two main phases: training of ensemble members, and synthesis (or fine-tuning) of the ensemble members into a policy that works in a new environment. The first phase of the algorithm involves training regular policy gradient or actor-critic agents in parallel but adding a term to the loss that encourages these policies to differ from each other. This causes the individual unimodal agents to explore the space of optimal policies and capture more of the multimodality of the environment than a single actor could. The second phase of DEFT involves synthesizing the component policies into a new policy that works well in a modified environment in one of two ways. To evaluate the performance of DEFT, we start with a base version of the Proximal Policy Optimization (PPO) algorithm and extend it with the modifications for DEFT. Our results show that the pretraining phase is effective in producing diverse policies in multimodal environments. DEFT often converges to a high reward significantly faster than alternatives, such as random initialization without DEFT and fine-tuning of ensemble members. While there is certainly more work to be done to analyze DEFT theoretically and extend it to be even more robust, we believe it provides a strong framework for capturing multimodality in environments while still using RL methods with simple policy representations.

📄 PDF Abstract BibTeX arXiv:2209.12412

Code (0)

등록된 구현이 없습니다.

Tasks

Ensemble Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Deft Scheduling of Dynamic Cloud Workflows with Varying Deadlines via Mixture-of-Experts

2026-05-31 · Ya Shen, Gang Chen, Hui Ma, Mengjie Zhang arxiv

Workflow scheduling in cloud computing demands the intelligent allocation of dynamically arriving, graph-structured workflows with varying deadlines onto ever-changing virtual machine resources. However, existing deep re…

Reinforcement Learning

DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer

2025-05-21 · Sona Elza Simon, Preethi Jyothi

Effective cross-lingual transfer remains a critical challenge in scaling the benefits of large language models from high-resource to low-resource languages. Towards this goal, prior studies have explored many approaches …

Cross-Lingual TransferNatural Language InferenceSentiment AnalysisSentiment Classification+1

DeftPunk at SemEval-2020 Task 6: Using RNN-ensemble for the Sentence Classification.

2020-12-01 · SEMEVAL 2020 · Jekaterina Kaparina, Anna Soboleva

This paper describes participation in DeftEval 2020 (part of SemEval sharing task competition), and is focused on the sentence classification. Our approach to the task was to create an ensemble of several RNNs combined w…

SentenceSentence Classification

MotherNets: Rapid Deep Ensemble Learning

2018-09-12 · Abdul Wasay, Brian Hentschel, Yuze Liao, Sanyuan Chen 외

Ensembles of deep neural networks significantly improve generalization accuracy. However, training neural network ensembles requires a large amount of computational resources and time. State-of-the-art approaches either …

ClusteringClustering EnsembleEnsemble LearningKnowledge Distillation

DEFT-VTON: Efficient Virtual Try-On with Consistent Generalised H-Transform

2025-09-16 · Xingzi Xu, Qi Li, Shuwen Qiu, Julien Han 외 arxiv

Diffusion models enable high-quality virtual try-on (VTO) with their established image synthesis abilities. Despite the extensive end-to-end training of large pre-trained models involved in current VTO methods, real-worl…

parameter-efficient fine-tuningVirtual Try-on