paper-with-me

Papers

RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning

2026-09-02 · Howard Qian, Yiting Chen, Yunfei Xie, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen, Chen Wei, Kaiyu Hang hf

Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains expensive and poorly suited to covering the long tail of real-world tasks. To address this bottleneck, we introduce RoboTok, an internet-scale data engine that, given a query human manipulation video, retrieves manipulation-relevant human demonstrations from web videos for training dexterous robot policies. Specifically, we learn a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames. This representation enables manipulation behaviors to be compared across variations in camera viewpoint, scene appearance, and actor occlusions, while remaining compact enough for efficient search and continual indexing over internet-scale video collections. We evaluate RoboTok against existing robot-data retrieval approaches on retrieval benchmarks and downstream robot policy performance. Our results show that RoboTok retrieves more relevant manipulation demonstrations and improves downstream task success, establishing hand-pose trajectory-aware retrieval as a way to make web video a scalable and continuously growing source of supervision for robot learning.

📄 PDF Abstract BibTeX arXiv:2609.03199

Code (4)

Rice-RobotPI-Lab/RoboTok-Code ★ 8
Tavish9/awesome-daily-AI-arxiv ★ 114
arxivsub/arXivSub_daily_arxiv ★ 4
🤗 Rice-RobotPI-Lab/robotok-public ★ 1

Similar Papers 제목 키워드 기반

Data Management Challenges for Internet-scale 3D Search Engines

2022-09-08 · James Williams, Shane Scott, Sean Wedig, Timur Hindanov 외

This paper describes the most significant data-related challenges involved in building internet-scale 3D search engines. The discussion centers on the most pressing data management issues in this domain, including model …

Information RetrievalManagementRetrieval

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning

2026-06-16 · Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen 외 arxiv

Internet videos constitute the largest reservoir of embodied human manipulation knowledge, yet converting arbitrary RGB footage into actionable robot training data remains a major bottleneck. Existing lab- or factory-col…

Learning about Canonical Views from Internet Image Collections

2012-12-01 · NeurIPS 2012 12 · Elad Mezuman, Yair Weiss

Although human object recognition is supposedly robust to viewpoint, much research on human perception indicates that there is a preferred or “canonical” view of objects. This phenomenon was discovered more than 30 years…

Object Recognition

Who Provides the Largest Megaphone? The Role of Google News in Promoting Russian State-Affiliated News Sources

2023-07-19 · Keeley Erhardt, Saurabh Khanna

The Internet has not only digitized but also democratized information access across the globe. This gradual but path-breaking move to online information propagation has resulted in search engines playing an increasingly …

Distilling Internet-Scale Vision-Language Models into Embodied Agents

2023-01-29 · Theodore Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 외

Instruction-following agents must ground language into their observation and action spaces. Learning to ground language is challenging, typically requiring domain-specific engineering or large quantities of human interac…

Instruction Following