paper-with-me

Papers

PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics

2026-04-26 · Tianyidan Xie, Zhentao Huang, Mingjie Wang, Xin Huang, Jun Zhou, Minglun Gong, Zili Yi arxiv

Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated physics simulators, they remain confined to 2D planar motions and fail to capture depth-aware spatial interactions. We introduce PhysLayer, a novel framework enabling language-guided, depth-aware layered animation of static images. PhysLayer consists of three key components: First, a language-guided scene understanding module that utilizes vision foundation models to decompose scenes into depth-based layers by analyzing object composition, material properties, and physical parameters. Second, a depth-aware layered physics simulation that extends 2D rigid-body dynamics with depth motion and perspective-consistent scaling, enabling more realistic object interactions without requiring full 3D reconstruction. Third, a physics-guided video synthesis module that integrates simulated trajectories with scene-aware relighting for temporally coherent results. Experimental results demonstrate improvements in CLIP-Similarity (+2.2\%), FID score (+9.3\%), and Motion-FID (+3\%), with human evaluation showing enhanced physical plausibility (+24\%) and text-video alignment (+35\%). Our approach provides a practical balance between physical realism and computational efficiency for controllable image animation.

📄 PDF Abstract BibTeX arXiv:2604.23574

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyScene Understanding3D ReconstructionVideo Generation

Similar Papers 제목 키워드 기반

3D Cinemagraphy from a Single Image

2023-03-10 · CVPR 2023 1 · Xingyi Li, Zhiguo Cao, Huiqiang Sun, Jianming Zhang 외

We present 3D Cinemagraphy, a new technique that marries 2D image animation with 3D photography. Given a single still image as input, our goal is to generate a video that contains both visual content animation and camera…

Image AnimationMotion Estimation

MagicAvatar: Multimodal Avatar Generation and Animation

2023-08-28 · Jianfeng Zhang, Hanshu Yan, Zhongcong Xu, Jiashi Feng 외

This report presents MagicAvatar, a framework for multimodal video generation and animation of human avatars. Unlike most existing methods that generate avatar-centric videos directly from multimodal inputs (e.g., text p…

Video Generation

Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling

2026-02-22 · Qi Sun, Can Wang, Jiaxiang Shang, Yingchun Liu 외 arxiv

Current 3D human animation methods struggle to achieve photorealism: kinematics-based approaches lack non-rigid dynamics (e.g., clothing dynamics), while methods that leverage video diffusion priors can synthesize non-ri…

Towards Multi-Layered 3D Garments Animation

2023-05-17 · ICCV 2023 1 · Yidi Shao, Chen Change Loy, Bo Dai

Mimicking realistic dynamics in 3D garment animations is a challenging task due to the complex nature of multi-layered garments and the variety of outer forces involved. Existing approaches mostly focus on single-layered…

Layered Depth Refinement with Mask Guidance

2022-06-07 · CVPR 2022 1 · Soo Ye Kim, Jianming Zhang, Simon Niklaus, Yifei Fan 외

Depth maps are used in a wide range of applications from 3D rendering to 2D image effects such as Bokeh. However, those predicted by single image depth estimation (SIDE) models often fail to capture isolated holes in obj…

Depth EstimationDepth PredictionImage MattingSelf-Supervised Learning