paper-with-me

홈 › Papers

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

2025-10-09 · Xiuwei Xu, Angyuan Ma, Hankun Li, Bingyao Yu, Zheng Zhu, Jie Zhou, Jiwen Lu arxiv

Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to work robustly under different spatial distribution of objects, environment and agent itself. To achieve this, substantial human demonstrations need to be collected to cover different spatial configurations for training a generalized visuomotor policy via imitation learning. Prior works explore a promising direction that leverages data generation to acquire abundant spatially diverse data from minimal source demonstrations. However, most approaches face significant sim-to-real gap and are often limited to constrained settings, such as fixed-base scenarios and predefined camera viewpoints. In this paper, we propose a real-to-real 3D data generation framework (R2RGen) that directly augments the pointcloud observation-action pairs to generate real-world data. R2RGen is simulator- and rendering-free, thus being efficient and plug-and-play. Specifically, we propose a unified three-stage framework, which (1) pre-processes source demonstrations under different camera setups in a shared 3D space with scene / trajectory parsing; (2) augments objects and robot's position with a group-wise backtracking strategy; (3) aligns the distribution of generated data with real-world 3D sensor using camera-aware post-processing. Empirically, R2RGen substantially enhances data efficiency on extensive experiments and demonstrates strong potential for scaling and application on mobile manipulation.

📄 PDF Abstract BibTeX arXiv:2510.08547

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spatially Grounded Long-Horizon Task Planning in the Wild

2026-03-13 · Sehun Jung, HyunJee Song, Dong-Hee Kim, Reuben Tan 외 arxiv

Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action plans expressed in natural language that gu…

Robot Manipulation

From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction Models

2025-01-01 · CVPR 2025 1 · German Barquero, Nadine Bertsch, Manojkumar Marramreddy, Carlos Chacón 외

In extended reality (XR), generating full-body motion of the users is important to understand their actions, drive their virtual avatars for social interaction, and convey a realistic sense of presence. While prior w…

FrictionMotion Generation

MetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative Modeling

2026-04-19 · Jinqi Cao, Zhiping Yu, Baihong Lin, Chenyang Liu 외 arxiv

Recent generative AI models have achieved remarkable breakthroughs in language and visual understanding. However, although these models can generate realistic visual content, their spatial scale remains confined to bound…

3D Generation

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

2026-04-21 · Xiangyang Luo, Xiaozhe Xin, Tao Feng, Xu Guo 외 arxiv

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capabilit…

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

2026-04-21 · Gene Chou, Charles Herrmann, Kyle Genova, Boyang Deng 외 arxiv

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consisten…

Autonomous DrivingVideo Generation