paper-with-me

홈 › Papers

HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion

2023-11-28 · CVPR 2024 1 · Jingbo Zhang, Xiaoyu Li, Qi Zhang, YanPei Cao, Ying Shan, Jing Liao

Generating a 3D human model from a single reference image is challenging because it requires inferring textures and geometries in invisible views while maintaining consistency with the reference image. Previous methods utilizing 3D generative models are limited by the availability of 3D training data. Optimization-based methods that lift text-to-image diffusion models to 3D generation often fail to preserve the texture details of the reference image, resulting in inconsistent appearances in different views. In this paper, we propose HumanRef, a 3D human generation framework from a single-view input. To ensure the generated 3D model is photorealistic and consistent with the input image, HumanRef introduces a novel method called reference-guided score distillation sampling (Ref-SDS), which effectively incorporates image guidance into the generation process. Furthermore, we introduce region-aware attention to Ref-SDS, ensuring accurate correspondence between different body regions. Experimental results demonstrate that HumanRef outperforms state-of-the-art methods in generating 3D clothed humans with fine geometry, photorealistic textures, and view-consistent appearances.

📄 PDF Abstract BibTeX arXiv:2311.16961

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationImage to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance

2024-07-09 · Guian Fang, Wenbiao Yan, Yuanfan Guo, Jianhua Han 외

Text-to-image diffusion models have significantly advanced in conditional image generation. However, these models usually struggle with accurately rendering images featuring humans, resulting in distorted limbs and other…

BenchmarkingConditional Image GenerationDiagnosticImage Generation+2

Referring to Any Person

2025-03-11 · Qing Jiang, Lin Wu, Zhaoyang Zeng, Tianhe Ren 외

Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referring to any person, holds substantial pra…

Large Language ModelMultimodal Large Language Modelobject-detectionObject Detection

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

2025-06-04 · Qing Jiang, Xingyu Chen, Zhaoyang Zeng, Junzhi Yu 외

Object referring aims to detect all objects in an image that match a given natural language description. We argue that a robust object referring model should be grounded, meaning its predictions should be both explainabl…

ObjectReferring Expression

Learning Multi-dimensional Human Preference for Text-to-Image Generation

2024-05-23 · CVPR 2024 1 · Sixian Zhang, Bohan Wang, Junqiang Wu, Yan Li 외

Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated imag…

Image GenerationText to Image GenerationText-to-Image Generation

FaithFill: Faithful Inpainting for Object Completion Using a Single Reference Image

2024-06-12 · Rupayan Mallick, Amr Abdalla, Sarah Adel Bargal

We present FaithFill, a diffusion-based inpainting object completion approach for realistic generation of missing object parts. Typically, multiple reference images are needed to achieve such realistic generation, otherw…

Object