paper-with-me

홈 › Papers

Learning Skills from Action-Free Videos

2025-12-23 · Hung-Chieh Fang, Kuo-Han Hung, Chu-Rong Chen, Po-Jung Chou, Chun-Kai Yang, Po-Chen Ko, Yu-Chiang Wang, Yueh-Hua Wu, Min-Hung Chen, Shao-Hua Sun arxiv

Learning from videos offers a promising path toward generalist robots by providing rich visual and temporal priors beyond what real robot datasets contain. While existing video generative models produce impressive visual predictions, they are difficult to translate into low-level actions. Conversely, latent-action models better align videos with actions, but they typically operate at the single-step level and lack high-level planning capabilities. We bridge this gap by introducing Skill Abstraction from Optical Flow (SOF), a framework that learns latent skills from large collections of action-free videos. Our key idea is to learn a latent skill space through an intermediate representation based on optical flow that captures motion information aligned with both video dynamics and robot actions. By learning skills in this flow-based latent space, SOF enables high-level planning over video-derived skills and allows for easier translation of these skills into actions. Experiments show that our approach consistently improves performance in both multitask and long-horizon settings, demonstrating the ability to acquire and compose skills directly from raw visual data.

📄 PDF Abstract BibTeX arXiv:2512.20052

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos

2025-12-10 · Hai Ci, Xiaokang Liu, Pei Yang, Yiren Song 외 arxiv

Robots that learn manipulation skills from everyday human videos could acquire broad capabilities without tedious robot data collection. We propose a video-to-video translation framework that converts ordinary human-obje…

Robot Manipulation

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos

2024-12-14 · Xin Liu, Yaran Chen

Current advanced policy learning methodologies have demonstrated the ability to develop expert-level strategies when provided enough information. However, their requirements, including task-specific rewards, expert-label…

Open-World Skill Discovery from Unsegmented Demonstrations

2025-03-11 · Jingwen Deng, ZiHao Wang, Shaofei Cai, Anji Liu 외

Learning skills in open-world environments is essential for developing agents capable of handling a variety of tasks by combining basic skills. Online demonstration videos are typically long but unsegmented, making them …

Boundary DetectionEvent SegmentationInstruction FollowingMinecraft+3

SportSkills: Physical Skill Learning from Sports Instructional Videos

2026-03-26 · Kumar Ashutosh, Chi Hsuan Wu, Kristen Grauman arxiv

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports …

Representation LearningVideo Retrieval

HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos

2025-09-20 · Haoyang Weng, Yitang Li, Nikhil Sobanbabu, Zihan Wang 외 arxiv

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framewo…

Reinforcement Learning