paper-with-me

홈 › Papers

RoboKoop: Efficient Control Conditioned Representations from Visual Input in Robotics using Koopman Operator

2024-09-04 · Hemant Kumawat, Biswadeep Chakraborty, Saibal Mukhopadhyay

Developing agents that can perform complex control tasks from high-dimensional observations is a core ability of autonomous agents that requires underlying robust task control policies and adapting the underlying visual representations to the task. Most existing policies need a lot of training samples and treat this problem from the lens of two-stage learning with a controller learned on top of pre-trained vision models. We approach this problem from the lens of Koopman theory and learn visual representations from robotic agents conditioned on specific downstream tasks in the context of learning stabilizing control for the agent. We introduce a Contrastive Spectral Koopman Embedding network that allows us to learn efficient linearized visual representations from the agent's visual data in a high dimensional latent space and utilizes reinforcement learning to perform off-policy control on top of the extracted representations with a linear controller. Our method enhances stability and control in gradient dynamics over time, significantly outperforming existing approaches by improving efficiency and accuracy in learning task policies over extended horizons.

📄 PDF Abstract BibTeX arXiv:2409.03107

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MichiGAN: Multi-Input-Conditioned Hair Image Generation for Portrait Editing

2020-10-30 · Zhentao Tan, Menglei Chai, Dongdong Chen, Jing Liao 외

Despite the recent success of face image generation with GANs, conditional hair editing remains challenging due to the under-explored complexity of its geometry and appearance. In this paper, we present MichiGAN (Multi-I…

Conditional Image GenerationImage Generation

Img2CAD: Conditioned 3D CAD Model Generation from Single Image with Structured Visual Geometry

2024-10-04 · Tianrun Chen, Chunan Yu, Yuanqi Hu, Jing Li 외

In this paper, we propose Img2CAD, the first approach to our knowledge that uses 2D image inputs to generate CAD models with editable parameters. Unlike existing AI methods for 3D model generation using text or image inp…

3D Reconstruction

Transforming Image Generation from Scene Graphs

2022-07-01 · Renato Sortino, Simone Palazzo, Concetto Spampinato

Generating images from semantic visual knowledge is a challenging task, that can be useful to condition the synthesis process in complex, subtle, and unambiguous ways, compared to alternatives such as class labels or tex…

DecoderImage GenerationImage Generation from Scene Graphs

PVI: Plug-in Visual Injection for Vision-Language-Action Models

2026-03-13 · Zezhou Zhang, Songxin Zhang, Xiao Xiong, Junjie Zhang 외 arxiv

VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for semantic abstraction and typically condi…

Language-Driven Representation Learning for Robotics

2023-02-24 · Siddharth Karamcheti, Suraj Nair, Annie S. Chen, Thomas Kollar 외

Recent work in visual representation learning for robotics demonstrates the viability of learning from large video datasets of humans performing everyday tasks. Leveraging methods such as masked autoencoding and contrast…

Contrastive LearningImitation LearningRepresentation LearningText Generation