paper-with-me

홈 › Papers

Synthesizing Moving People with 3D Control

2024-01-19 · Boyi Li, Junming Chen, Jathushan Rajasegaran, Yossi Gandelsman, Alexei A. Efros, Jitendra Malik

In this paper, we present a diffusion model-based framework for animating people from a single image for a given target 3D motion sequence. Our approach has two core components: a) learning priors about invisible parts of the human body and clothing, and b) rendering novel body poses with proper clothing and texture. For the first part, we learn an in-filling diffusion model to hallucinate unseen parts of a person given a single image. We train this model on texture map space, which makes it more sample-efficient since it is invariant to pose and viewpoint. Second, we develop a diffusion-based rendering pipeline, which is controlled by 3D human poses. This produces realistic renderings of novel poses of the person, including clothing, hair, and plausible in-filling of unseen regions. This disentangled approach allows our method to generate a sequence of images that are faithful to the target motion in the 3D pose and, to the input image in terms of visual similarity. In addition to that, the 3D control allows various synthetic camera trajectories to render a person. Our experiments show that our method is resilient in generating prolonged motions and varied challenging and complex poses compared to prior methods. Please check our website for more details: https://boyiliee.github.io/3DHM.github.io/.

📄 PDF Abstract BibTeX arXiv:2401.10889

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular Video

2022-01-11 · CVPR 2022 1 · Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron 외

We introduce a free-viewpoint rendering method -- HumanNeRF -- that works on a given monocular video of a human performing complex body motions, e.g. a video from YouTube. Our method enables pausing the video at any fram…

View-LSTM: Novel-View Video Synthesis Through View Decomposition

2019-10-01 · ICCV 2019 10 · Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro

We tackle the problem of synthesizing a video of multiple moving people as seen from a novel view, given only an input video and depth information or human poses of the novel view as prior. This problem requires a model …

Action Recognition

Removing Objects From Neural Radiance Fields

2022-12-22 · CVPR 2023 1 · Silvan Weder, Guillermo Garcia-Hernando, Aron Monszpart, Marc Pollefeys 외

Neural Radiance Fields (NeRFs) are emerging as a ubiquitous scene representation that allows for novel view synthesis. Increasingly, NeRFs will be shareable with other people. Before sharing a NeRF, though, it might be d…

Image InpaintingNeRFNovel View Synthesis

The Un-Kidnappable Robot: Acoustic Localization of Sneaking People

2023-10-05 · Mengyu Yang, Patrick Grady, Samarth Brahmbhatt, Arun Balajee Vasudevan 외

How easy is it to sneak up on a robot? We examine whether we can detect people using only the incidental sounds they produce as they move, even when they try to be quiet. We collect a robotic dataset of high-quality 4-ch…

Learning the Depths of Moving People by Watching Frozen People

2019-04-25 · CVPR 2019 6 · Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker 외

We present a method for predicting dense depth in scenarios where both a monocular camera and people in the scene are freely moving. Existing methods for recovering depth for dynamic, non-rigid objects from monocular vid…

Depth EstimationDepth Prediction