paper-with-me

홈 › Papers

DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation

2024-04-09 · Junkai Yan, Yipeng Gao, Qize Yang, Xihan Wei, Xuansong Xie, AnCong Wu, Wei-Shi Zheng

Text-to-3D generation, which synthesizes 3D assets according to an overall text description, has significantly progressed. However, a challenge arises when the specific appearances need customizing at designated viewpoints but referring solely to the overall description for generating 3D objects. For instance, ambiguity easily occurs when producing a T-shirt with distinct patterns on its front and back using a single overall text guidance. In this work, we propose DreamView, a text-to-image approach enabling multi-view customization while maintaining overall consistency by adaptively injecting the view-specific and overall text guidance through a collaborative text guidance injection module, which can also be lifted to 3D generation via score distillation sampling. DreamView is trained with large-scale rendered multi-view images and their corresponding view-specific texts to learn to balance the separate content manipulation in each view and the global consistency of the overall object, resulting in a dual achievement of customization and consistency. Consequently, DreamView empowers artists to design 3D objects creatively, fostering the creation of more innovative and diverse 3D assets. Code and model will be released at https://github.com/iSEE-Laboratory/DreamView.

📄 PDF Abstract BibTeX arXiv:2404.06119

Code (1)

isee-laboratory/dreamview 공식 구현 pytorch

Tasks

3D GenerationText to 3D

Similar Papers 제목 키워드 기반

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

2026-04-19 · Hong Jiang, Wensong Song, Zongxin Yang, Ruijie Quan 외 arxiv

Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmen…

Image EditingPoint Clouds

Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference

2026-05-24 · Agata Żywot, Iason Skylitsis, Thijmen Nijdam, Zoe Tzifa-Kratira 외 arxiv

Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either requi…

Style Transfer

Face Swap via Diffusion Model

2024-03-02 · Feifei Wang

This technical report presents a diffusion model based framework for face swapping between two portrait images. The basic framework consists of three components, i.e., IP-Adapter, ControlNet, and Stable Diffusion's inpai…

Face AlignmentFace DetectionFace SwappingFacial Inpainting+1

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration

2026-06-23 · Yang Zhou, Wenxue Li, Peng Zhang, Yifei Chen 외 arxiv

Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address thr…

Video RestorationVideo Generation

Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention

2025-08-04 · Kyungmin Jo, Jooyeol Yun, Jaegul Choo arxiv

While large-scale text-to-image diffusion models enable the generation of high-quality, diverse images from text prompts, these prompts struggle to capture intricate details, such as textures, preventing the user intent …

Image Generation