paper-with-me

홈 › Papers

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation

2026-04-09 · Jingjing Wang, Zhengdong Hong, Chong Bao, Yuke Zhu, Junhan Sun, Guofeng Zhang arxiv

Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforcement learning, imitation learning, and vision-language-action-models (VLAs), often struggle with novel tasks and unseen environments. Another promising direction is to explore generalizable representations that capture fine-grained spatial and geometric relations for open-world manipulation. While large-language-model (LLMs) and vision-language-model (VLMs) provide strong semantic reasoning based on language or annotated 2D representations, their limited 3D awareness restricts their applicability to fine-grained manipulation. To address this, we propose LAMP, which lifts image-editing as 3D priors to extract inter-object 3D transformations as continuous, geometry-aware representations. Our key insight is that image-editing inherently encodes rich 2D spatial cues, and lifting these implicit cues into 3D transformations provides fine-grained and accurate guidance for open-world manipulation. Extensive experiments demonstrate that \codename delivers precise 3D transformations and achieves strong zero-shot generalization in open-world manipulation. Project page: https://zju3dv.github.io/LAMP/.

📄 PDF Abstract BibTeX arXiv:2604.08475

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationReinforcement Learning

Similar Papers 제목 키워드 기반

VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors

2025-03-03 · CVPR 2025 1 · Juil Koo, Paul Guerrero, Chun-Hao Paul Huang, Duygu Ceylan 외

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have …

3D ReconstructionObjectVideo Editing

AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing

2025-12-08 · Ziming Hong, Tianyu Huang, Runnan Chen, Shanshan Ye 외 arxiv

Recent studies have extended diffusion-based instruction-driven 2D image editing pipelines to 3D Gaussian Splatting (3DGS), enabling faithful manipulation of 3DGS assets and greatly advancing 3DGS content creation. Howev…

Image Editing

Generative Image as Action Models

2024-07-10 · Mohit Shridhar, Yat Long Lo, Stephen James

Image-generation diffusion models have been fine-tuned to unlock new capabilities such as image-editing and novel view synthesis. Can we similarly unlock image-generation models for visuomotor control? We present GENIMA,…

Image GenerationRobot ManipulationRobot Manipulation Generalization

C3Editor: Achieving Controllable Consistency in 2D Model for 3D Editing

2025-10-06 · Zeng Tao, Zheng Ding, Zeyuan Chen, Xiang Zhang 외 arxiv

Existing 2D-lifting-based 3D editing methods often encounter challenges related to inconsistency, stemming from the lack of view-consistent 2D editing models and the difficulty of ensuring consistent editing across multi…

LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

2024-12-12 · Yabo Chen, Chen Yang, Jiemin Fang, Xiaopeng Zhang 외

Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer pr…

3D ReconstructionImage to 3DVideo Generation