paper-with-me

Papers

Physically-based Lighting Generation for Robotic Manipulation

2025-08-02 · Shutong Jin, Lezhong Wang, Ben Temming, Florian T. Pokorny arxiv

In this paper, we propose the first framework that leverages physically-based inverse rendering for novel lighting generation on existing real-world human demonstrations of robotic manipulation tasks. Specifically, inverse rendering decomposes the first frame in each demonstration into geometric (surface normal, depth) and material (albedo, roughness, metallic) properties, which are then used to render appearance changes under different lighting sources. To improve efficiency and maintain consistency across each generated sequence, we fine-tune Stable Video Diffusion on robot execution videos for temporal lighting propagation. We evaluate our framework by measuring the visual quality of the generated sequences, assessing its effectiveness in improving the imitation learning policy performance (38.75\%) under six unseen real-world lighting conditions, and conduct ablation studies on individual modules of the proposed framework. We further showcase three downstream applications enabled by the proposed framework: background generation, object texture generation and distractor positioning. The code for the framework will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2508.01442

Code (0)

등록된 구현이 없습니다.

Tasks

Inverse Rendering

Similar Papers 제목 키워드 기반

Robot Learning from a Physical World Model

2025-11-10 · Jiageng Mao, Sicheng He, Hao-Ning Wu, Yang You 외 arxiv

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language comm…

Reinforcement LearningVideo Generation

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

2026-04-21 · Feng Jiang, Yang Chen, Kyle Xu, Yuchen Liu 외 arxiv

Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied ma…

Spatial Reasoning

Physically Grounded Vision-Language Models for Robotic Manipulation

2023-09-05 · Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao 외

Recent advances in vision-language models (VLMs) have led to improved performance on tasks such as visual question answering and image captioning. Consequently, these models are now well-positioned to reason about the ph…

Image CaptioningLanguage ModellingLarge Language ModelObject+2

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

2026-02-18 · Zijian Song, Qichang Li, Sihan Qin, Yuhao Chen 외 arxiv

The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning Physics from Pretrained Video Generation…

Physical IntuitionVideo Generation

High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting

2025-10-12 · Haoyu Zhao, Cheng Zeng, Linghao Zhuang, Yaxi Zhao 외 arxiv

The scalability of robotic learning is fundamentally bottlenecked by the significant cost and labor of real-world data collection. While simulated data offers a scalable alternative, it often fails to generalize to the r…