paper-with-me

홈 › Papers

Pre-trained Visual Representations Generalize Where it Matters in Model-Based Reinforcement Learning

2025-09-16 · Scott Jones, Liyou Zhou, Sebastian W. Pattinson arxiv

In visuomotor policy learning, the control policy for the robotic agent is derived directly from visual inputs. The typical approach, where a policy and vision encoder are trained jointly from scratch, generalizes poorly to novel visual scene changes. Using pre-trained vision models (PVMs) to inform a policy network improves robustness in model-free reinforcement learning (MFRL). Recent developments in Model-based reinforcement learning (MBRL) suggest that MBRL is more sample-efficient than MFRL. However, counterintuitively, existing work has found PVMs to be ineffective in MBRL. Here, we investigate PVM's effectiveness in MBRL, specifically on generalization under visual domain shifts. We show that, in scenarios with severe shifts, PVMs perform much better than a baseline model trained from scratch. We further investigate the effects of varying levels of fine-tuning of PVMs. Our results show that partial fine-tuning can maintain the highest average task performance under the most extreme distribution shifts. Our results demonstrate that PVMs are highly successful in promoting robustness in visual policy learning, providing compelling evidence for their wider adoption in model-based robotic learning applications.

📄 PDF Abstract BibTeX arXiv:2509.12531

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

What Matters to You? Towards Visual Representation Alignment for Robot Learning

2023-10-11 · Ran Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik 외

When operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs like RGB images, their rewards will inevitably use visual representa…

Zero-shot Generalization

Viewpoint Matters: Dynamically Optimizing Viewpoints with Masked Autoencoder for Visual Manipulation

2026-02-04 · Pengfei Yi, Yifan Han, Junyan Li, Litao Liu 외 arxiv

Robotic manipulation continues to be a challenge, and imitation learning (IL) enables robots to learn tasks from expert demonstrations. Current IL methods typically rely on fixed camera setups, where cameras are manually…

Task Formulation Matters When Learning Continually: A Case Study in Visual Question Answering

2022-09-30 · Mavina Nikandrou, Lu Yu, Alessandro Suglia, Ioannis Konstas 외

Continual learning aims to train a model incrementally on a sequence of tasks without forgetting previous knowledge. Although continual learning has been widely studied in computer vision, its application to Vision+Langu…

Continual LearningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Multi-modal Cycle-consistent Generalized Zero-Shot Learning

2018-08-01 · ECCV 2018 9 · Rafael Felix, B. G. Vijay Kumar, Ian Reid, Gustavo Carneiro

In generalized zero shot learning (GZSL), the set of classes are split into seen and unseen classes, where training relies on the semantic features of the seen and unseen classes and the visual representations of only th…

General ClassificationGeneralized Zero-Shot LearningZero-Shot Learning

Focus On What Matters: Separated Models For Visual-Based RL Generalization

2024-09-29 · Di Zhang, Bowen Lv, Hai Zhang, Feifan Yang 외

A primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliary tasks to enhance generalization, few a…

Image ReconstructionReinforcement Learning (RL)Representation Learning