paper-with-me

Papers

Human-oriented Representation Learning for Robotic Manipulation

2023-10-04 · Mingxiao Huo, Mingyu Ding, Chenfeng Xu, Thomas Tian, Xinghao Zhu, Yao Mu, Lingfeng Sun, Masayoshi Tomizuka, Wei Zhan

Humans inherently possess generalizable visual representations that empower them to efficiently explore and interact with the environments in manipulation tasks. We advocate that such a representation automatically arises from simultaneously learning about multiple simple perceptual skills that are critical for everyday scenarios (e.g., hand detection, state estimate, etc.) and is better suited for learning robot manipulation policies compared to current state-of-the-art visual representations purely based on self-supervised objectives. We formalize this idea through the lens of human-oriented multi-task fine-tuning on top of pre-trained visual encoders, where each task is a perceptual skill tied to human-environment interactions. We introduce Task Fusion Decoder as a plug-and-play embedding translator that utilizes the underlying relationships among these perceptual skills to guide the representation learning towards encoding meaningful structure for what's important for all perceptual skills, ultimately empowering learning of downstream robotic manipulation tasks. Extensive experiments across a range of robotic tasks and embodiments, in both simulations and real-world environments, show that our Task Fusion Decoder consistently improves the representation of three state-of-the-art visual encoders including R3M, MVP, and EgoVLP, for downstream manipulation policy-learning. Project page: https://sites.google.com/view/human-oriented-robot-learning

📄 PDF Abstract BibTeX arXiv:2310.03023

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderHand DetectionRepresentation LearningRobot Manipulation

Similar Papers 제목 키워드 기반

SoftGPT: Learn Goal-oriented Soft Object Manipulation Skills by Generative Pre-trained Heterogeneous Graph Transformer

2023-06-22 · Junjia Liu, Zhihao LI, WanYu Lin, Sylvain Calinon 외

Soft object manipulation tasks in domestic scenes pose a significant challenge for existing robotic skill learning techniques due to their complex dynamics and variable shape characteristics. Since learning new manipulat…

Object

ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation

2025-09-25 · Dekun Lu, Wei Gao, Kui Jia arxiv

End-to-end robot manipulation policies offer significant potential for enabling embodied agents to understand and interact with the world. Unlike traditional modular pipelines, end-to-end learning mitigates key limitatio…

Robot Manipulation

A Road-map to Robot Task Execution with the Functional Object-Oriented Network

2021-06-01 · David Paulius, Alejandro Agostini, Yu Sun, Dongheui Lee

Following work on joint object-action representations, the functional object-oriented network (FOON) was introduced as a knowledge graph representation for robots. Taking the form of a bipartite graph, a FOON contains sy…

Task Planning

ToMPC: Task-oriented Model Predictive Control via ADMM for Safe Robotic Manipulation

2026-03-14 · Xinyu Jia, Wenxin Wang, Jun Yang, Yongping Pan 외 arxiv

This paper proposes a task-oriented model predictive control (ToMPC) framework for safe and efficient robotic manipulation in open workspaces. The framework unifies collision-free motion and robot-environment interaction…

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

2025-08-18 · Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang 외 arxiv

Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditional rule-based methods fail to scale or generalize in unstructured, novel environ…

Reinforcement Learning