paper-with-me

홈 › Papers

Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement Learning

2020-10-26 · NeurIPS 2020 12 · Younggyo Seo, Kimin Lee, Ignasi Clavera, Thanard Kurutach, Jinwoo Shin, Pieter Abbeel

Model-based reinforcement learning (RL) has shown great potential in various control tasks in terms of both sample-efficiency and final performance. However, learning a generalizable dynamics model robust to changes in dynamics remains a challenge since the target transition dynamics follow a multi-modal distribution. In this paper, we present a new model-based RL algorithm, coined trajectory-wise multiple choice learning, that learns a multi-headed dynamics model for dynamics generalization. The main idea is updating the most accurate prediction head to specialize each head in certain environments with similar dynamics, i.e., clustering environments. Moreover, we incorporate context learning, which encodes dynamics-specific information from past experiences into the context latent vector, enabling the model to perform online adaptation to unseen environments. Finally, to utilize the specialized prediction heads more effectively, we propose an adaptive planning method, which selects the most accurate prediction head over a recent experience. Our method exhibits superior zero-shot generalization performance across a variety of control tasks, compared to state-of-the-art RL methods. Source code and videos are available at https://sites.google.com/view/trajectory-mcl.

📄 PDF Abstract BibTeX arXiv:2010.13303

Code (1)

younggyoseo/trajectory_mcl 공식 구현 tf

Tasks

ClusteringModel-based Reinforcement LearningMultiple-choicePredictionreinforcement-learningReinforcement Learning (RL)Zero-shot Generalization

Similar Papers 제목 키워드 기반

Higher-Order Transformer Derivative Estimates for Explicit Pathwise Learning Guarantees

2024-05-26 · Yannick Limmer, Anastasis Kratsios, Xuwei Yang, Raeid Saqur 외

An inherent challenge in computing fully-explicit generalization bounds for transformers involves obtaining covering number estimates for the given transformer class $T$. Crude estimates rely on a uniform upper bound on …

Generalization Bounds

Detection- and Trajectory-Level Exclusion in Multiple Object Tracking

2013-06-01 · CVPR 2013 6 · Anton Milan, Konrad Schindler, Stefan Roth

When tracking multiple targets in crowded scenarios, modeling mutual exclusion between distinct targets becomes important at two levels: (1) in data association, each target observation should support at most one traject…

Multiple Object TrackingObjectObject Tracking

Temperature-driven inversion and nonlinear dynamics in ChatGPT-like AIs

2026-08-02 · Neil F. Johnson, Frank Yingjie Huo, Bella Xinrui Li arxiv

Increasing the temperature of an ordinary many-state system increases access to a wider range of states and hence increases its entropy. We find the opposite in ChatGPT-like AIs, even though raising the decoder temperatu…

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

2026-06-03 · Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi arxiv

In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function. Concretely, we consider a realizable setting where inputs are dra…

Longitudinal Flow Matching for Trajectory Modeling

2025-10-03 · Mohammad Mohaiminul Islam, Thijs P. Kuipers, Sharvaree Vadgama, Coen de Vente 외 arxiv

Generative models for sequential data often struggle with sparsely sampled and high-dimensional trajectories, typically reducing the learning of dynamics to pairwise transitions. We propose Interpolative Multi-Marginal F…

Trajectory Modeling