paper-with-me

Papers

Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

2024-05-02 · Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta, Shubham Tulsiani

We seek to learn a generalizable goal-conditioned policy that enables zero-shot robot manipulation: interacting with unseen objects in novel scenes without test-time adaptation. While typical approaches rely on a large amount of demonstration data for such generalization, we propose an approach that leverages web videos to predict plausible interaction plans and learns a task-agnostic transformation to obtain robot actions in the real world. Our framework,Track2Act predicts tracks of how points in an image should move in future time-steps based on a goal, and can be trained with diverse videos on the web including those of humans and robots manipulating everyday objects. We use these 2D track predictions to infer a sequence of rigid transforms of the object to be manipulated, and obtain robot end-effector poses that can be executed in an open-loop manner. We then refine this open-loop plan by predicting residual actions through a closed loop policy trained with a few embodiment-specific demonstrations. We show that this approach of combining scalably learned track prediction with a residual policy requiring minimal in-domain robot-specific data enables diverse generalizable robot manipulation, and present a wide array of real-world robot manipulation results across unseen tasks, objects, and scenes. https://homangab.github.io/track2act/

📄 PDF Abstract BibTeX arXiv:2405.01527

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationTest-time Adaptation

Similar Papers 제목 키워드 기반

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos

2026-03-09 · Adam Hung, Bardienus Pieter Duisterhof, Jeffrey Ichnowski arxiv

Learning manipulation policies from human videos could greatly reduce the need for expensive robot demonstrations, but existing approaches typically require restrictive assumptions such as choreographed human motions, pr…

Fast Encoder-Based 3D from Casual Videos via Point Track Processing

2024-04-10 · Yoni Kasten, Wuyue Lu, Haggai Maron

This paper addresses the long-standing challenge of reconstructing 3D structures from videos with dynamic content. Current approaches to this problem were not designed to operate on casual videos recorded by standard cam…

Point Tracking

DriveTrack: A Benchmark for Long-Range Point Tracking in Real-World Videos

2023-12-15 · CVPR 2024 1 · Arjun Balasingam, Joseph Chandler, Chenning Li, Zhoutong Zhang 외

This paper presents DriveTrack, a new benchmark and data generation framework for long-range keypoint tracking in real-world videos. DriveTrack is motivated by the observation that the accuracy of state-of-the-art tracke…

Autonomous DrivingPoint Tracking

Modality-Autoregressive World-Action Models

2026-09-15 · Adam Hung, Bardienus P. Duisterhof, Deva Ramanan, Jeffrey Ichnowski arxiv

World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more effici…

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

2026-07-19 · Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang 외 hf

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to comp…