paper-with-me

홈 › Papers

MVHuman: Tailoring 2D Diffusion with Multi-view Sampling For Realistic 3D Human Generation

2023-12-15 · Suyi Jiang, Haimin Luo, Haoran Jiang, Ziyu Wang, Jingyi Yu, Lan Xu

Recent months have witnessed rapid progress in 3D generation based on diffusion models. Most advances require fine-tuning existing 2D Stable Diffsuions into multi-view settings or tedious distilling operations and hence fall short of 3D human generation due to the lack of diverse 3D human datasets. We present an alternative scheme named MVHuman to generate human radiance fields from text guidance, with consistent multi-view images directly sampled from pre-trained Stable Diffsuions without any fine-tuning or distilling. Our core is a multi-view sampling strategy to tailor the denoising processes of the pre-trained network for generating consistent multi-view images. It encompasses view-consistent conditioning, replacing the original noises with ``consistency-guided noises'', optimizing latent codes, as well as utilizing cross-view attention layers. With the multi-view images through the sampling process, we adopt geometry refinement and 3D radiance field generation followed by a subsequent neural blending scheme for free-view rendering. Extensive experiments demonstrate the efficacy of our method, as well as its superiority to state-of-the-art 3D human generation methods.

📄 PDF Abstract BibTeX arXiv:2312.10120

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationDenoising

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization

2025-05-03 · Chenghong Li, Hongjie Liao, YiHao Zhi, Xihe Yang 외

In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while significant progress has been achieved …

MVHumanNet: A Large-scale Dataset of Multi-view Daily Dressing Human Captures

2023-12-05 · CVPR 2024 1 · Zhangyang Xiong, Chenghong Li, Kenkun Liu, Hongjie Liao 외

In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while remarkable progress has been made with …

Action RecognitionImage GenerationNeRF

PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human Modeling

2024-03-24 · CVPR 2024 1 · Xiaoyun Zheng, Liwei Liao, Xufeng Li, Jianbo Jiao 외

High-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable efforts invested in developing various capture systems a…

NeRFNovel View Synthesis

MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis

2025-10-08 · Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang 외 arxiv

Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current method…

Monocular Depth EstimationNovel View SynthesisVideo GenerationPoint Clouds

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

2026-08-20 · Yudong Jin, Tao Xie, Qihang Zhang, Zehong Shen 외 arxiv

We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existi…