paper-with-me

Papers

PhysGen3D: Crafting a Miniature Interactive World from a Single Image

2025-03-26 · CVPR 2025 1 · Boyuan Chen, Hanxiao Jiang, Shaowei Liu, Saurabh Gupta, Yunzhu Li, Hao Zhao, Shenlong Wang

Envisioning physically plausible outcomes from a single image requires a deep understanding of the world's dynamics. To address this, we introduce PhysGen3D, a novel framework that transforms a single image into an amodal, camera-centric, interactive 3D scene. By combining advanced image-based geometric and semantic understanding with physics-based simulation, PhysGen3D creates an interactive 3D world from a static image, enabling us to "imagine" and simulate future scenarios based on user input. At its core, PhysGen3D estimates 3D shapes, poses, physical and lighting properties of objects, thereby capturing essential physical attributes that drive realistic object interactions. This framework allows users to specify precise initial conditions, such as object speed or material properties, for enhanced control over generated video outcomes. We evaluate PhysGen3D's performance against closed-source state-of-the-art (SOTA) image-to-video models, including Pika, Kling, and Gen-3, showing PhysGen3D's capacity to generate videos with realistic physics while offering greater flexibility and fine-grained control. Our results show that PhysGen3D achieves a unique balance of photorealism, physical plausibility, and user-driven interactivity, opening new possibilities for generating dynamic, physics-grounded video from an image.

📄 PDF Abstract BibTeX arXiv:2503.20746

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

2026-02-18 · Zijian Song, Qichang Li, Sihan Qin, Yuhao Chen 외 arxiv

The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning Physics from Pretrained Video Generation…

Physical IntuitionVideo Generation

PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation

2024-09-27 · Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, Shenlong Wang

We present PhysGen, a novel image-to-video generation method that converts a single image and an input condition (e.g., force and torque applied to an object in the image) to produce a realistic, physically plausible, an…

Image to Video GenerationVideo Generation

Seeing A 3D World in A Grain of Sand

2025-03-01 · CVPR 2025 1 · Yufan Zhang, Yu Ji, Yu Guo, Jinwei Ye

We present a snapshot imaging technique for recovering 3D surrounding views of miniature scenes. Due to their intricacy, miniature scenes with objects sized in millimeters are difficult to reconstruct, yet miniatures are…

3DGS3D ReconstructionNovel View SynthesisSand

Vid2World: Crafting Video Diffusion Models to Interactive World Models

2025-05-20 · Siqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao 외

World models, which predict transitions based on history observation and action sequences, have shown great promise in improving data efficiency for sequential decision making. However, existing world models often requir…

Robot ManipulationSequential Decision Making

Efficient Cloth Simulation using Miniature Cloth and Upscaling Deep Neural Networks

2019-07-09 · Tae Min Lee, Young Jin Oh, In-Kwon Lee

Cloth simulation requires a fast and stable method for interactively and realistically visualizing fabric materials using computer graphics. We propose an efficient cloth simulation method using miniature cloth simulatio…