paper-with-me

Papers

ViSA-Flow: Accelerating Robot Skill Learning via Large-Scale Video Semantic Action Flow

2025-05-02 · Changhe Chen, Quantao Yang, Xiaohao Xu, Nima Fazeli, Olov Andersson

One of the central challenges preventing robots from acquiring complex manipulation skills is the prohibitive cost of collecting large-scale robot demonstrations. In contrast, humans are able to learn efficiently by watching others interact with their environment. To bridge this gap, we introduce semantic action flow as a core intermediate representation capturing the essential spatio-temporal manipulator-object interactions, invariant to superficial visual differences. We present ViSA-Flow, a framework that learns this representation self-supervised from unlabeled large-scale video data. First, a generative model is pre-trained on semantic action flows automatically extracted from large-scale human-object interaction video data, learning a robust prior over manipulation structure. Second, this prior is efficiently adapted to a target robot by fine-tuning on a small set of robot demonstrations processed through the same semantic abstraction pipeline. We demonstrate through extensive experiments on the CALVIN benchmark and real-world tasks that ViSA-Flow achieves state-of-the-art performance, particularly in low-data regimes, outperforming prior methods by effectively transferring knowledge from human video observation to robotic execution. Videos are available at https://visaflow-web.github.io/ViSAFLOW.

📄 PDF Abstract BibTeX arXiv:2505.01288

Code (0)

등록된 구현이 없습니다.

Tasks

Human-Object Interaction Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Advancing Improvisation in Human-Robot Construction Collaboration: Taxonomy and Research Roadmap

2026-01-23 · David Wireko Atibila, Vineet R. Kamat, Carol C. Menassa arxiv

The construction industry faces productivity stagnation, skilled labor shortages, and safety concerns. While robotic automation offers solutions, construction robots struggle to adapt to unstructured, dynamic sites. Cent…

SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization

2026-06-04 · Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Shusen Liu 외 arxiv

Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. While general-purpose coding agents show strong capabilities, they of…

Residual Skill Policies: Learning an Adaptable Skill-based Action Space for Reinforcement Learning for Robotics

2022-11-04 · Krishan Rana, Ming Xu, Brendan Tidd, Michael Milford 외

Skill-based reinforcement learning (RL) has emerged as a promising strategy to leverage prior knowledge for accelerated robot learning. Skills are typically extracted from expert demonstrations and are embedded into a la…

Reinforcement Learning (RL)

A Roadmap for Embodied and Social Grounding in LLMs

2024-09-25 · Sara Incao, Carlo Mazzola, Giulia Belgiovine, Alessandra Sciutti

The fusion of Large Language Models (LLMs) and robotic systems has led to a transformative paradigm in the robotic field, offering unparalleled capabilities not only in the communication domain but also in skills like mu…

Differentiable Skill Optimisation for Powder Manipulation in Laboratory Automation

2025-10-01 · Minglun Wei, Xintong Yang, Yu-Kun Lai, Ze Ji arxiv

Robotic automation is accelerating scientific discovery by reducing manual effort in laboratory workflows. However, precise manipulation of powders remains challenging, particularly in tasks such as transport that demand…

Reinforcement Learning