paper-with-me

홈 › Papers

WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation

2025-10-08 · Zezhong Qian, Xiaowei Chi, Yuming Li, Shizun Wang, Zhiyuan Qin, Xiaozhu Ju, Sirui Han, Shanghang Zhang arxiv

Wrist-view observations are crucial for VLA models as they capture fine-grained hand-object interactions that directly enhance manipulation performance. Yet large-scale datasets rarely include such recordings, resulting in a substantial gap between abundant anchor views and scarce wrist views. Existing world models cannot bridge this gap, as they require a wrist-view first frame and thus fail to generate wrist-view videos from anchor views alone. Amid this gap, recent visual geometry models such as VGGT emerge with geometric and cross-view priors that make it possible to address extreme viewpoint shifts. Inspired by these insights, we propose WristWorld, the first 4D world model that generates wrist-view videos solely from anchor views. WristWorld operates in two stages: (i) Reconstruction, which extends VGGT and incorporates our Spatial Projection Consistency (SPC) Loss to estimate geometrically consistent wrist-view poses and 4D point clouds; (ii) Generation, which employs our video generation model to synthesize temporally coherent wrist-view videos from the reconstructed perspective. Experiments on Droid, Calvin, and Franka Panda demonstrate state-of-the-art video generation with superior spatial consistency, while also improving VLA performance, raising the average task completion length on Calvin by 3.81% and closing 42.4% of the anchor-wrist view gap.

📄 PDF Abstract BibTeX arXiv:2510.07313

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationPoint Clouds

Similar Papers 제목 키워드 기반

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

2026-06-16 · Yuhang Huang, Xuan Lv, Junyan Xu, Zhiyuan Yu 외 arxiv

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on mult…

D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation

2025-05-08 · I-Chun Arthur Liu, Jason Chen, Gaurav Sukhatme, Daniel Seita

Learning bimanual manipulation is challenging due to its high dimensionality and tight coordination required between two arms. Eye-in-hand imitation learning, which uses wrist-mounted cameras, simplifies perception by fo…

Data AugmentationImitation Learning

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

2025-12-29 · Guo Ye, Zexi Zhang, Xu Zhao, Shang Wu 외 arxiv

Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain blind to physical contact. Consequently, they struggle with contact-rich manipul…

ByteWrist: A Parallel Robotic Wrist Enabling Flexible and Anthropomorphic Motion for Confined Spaces

2025-09-22 · Jiawen Tian, Liqun Huang, Zhongren Cui, Jingchao Qiao 외 arxiv

This paper introduces ByteWrist, a novel highly-flexible and anthropomorphic parallel wrist for robotic manipulation. ByteWrist addresses the critical limitations of existing serial and parallel wrists in narrow-space op…

DexWrist: A Robotic Wrist for Constrained and Dynamic Manipulation

2025-07-01 · Martin Peticco, Gabriella Ulloa, John Marangola, Nitish Dashora 외 arxiv

Development of dexterous manipulation hardware has primarily focused on hands and grippers. However, these end-effectors are often paired with bulky and highly stiff wrists that limit performance in human environments. M…