paper-with-me

홈 › Papers

DreamWalk: Style Space Exploration using Diffusion Guidance

2024-04-04 · Michelle Shu, Charles Herrmann, Richard Strong Bowen, Forrester Cole, Ramin Zabih

Text-conditioned diffusion models can generate impressive images, but fall short when it comes to fine-grained control. Unlike direct-editing tools like Photoshop, text conditioned models require the artist to perform "prompt engineering," constructing special text sentences to control the style or amount of a particular subject present in the output image. Our goal is to provide fine-grained control over the style and substance specified by the prompt, for example to adjust the intensity of styles in different regions of the image (Figure 1). Our approach is to decompose the text prompt into conceptual elements, and apply a separate guidance term for each element in a single diffusion process. We introduce guidance scale functions to control when in the diffusion process and \emph{where} in the image to intervene. Since the method is based solely on adjusting diffusion guidance, it does not require fine-tuning or manipulating the internal layers of the diffusion model's neural network, and can be used in conjunction with LoRA- or DreamBooth-trained models (Figure2). Project page: https://mshu1.github.io/dreamwalk.github.io/

📄 PDF Abstract BibTeX arXiv:2404.03145

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation

2023-08-14 · ICCV 2023 1 · Hanqing Wang, Wei Liang, Luc van Gool, Wenguan Wang

VLN-CE is a recently released embodied task, where AI agents need to navigate a freely traversable environment to reach a distant target location, given language instructions. It poses great challenges due to the huge sp…

Decision MakingNavigateVision-Language Navigation

Unifying Human Motion Synthesis and Style Transfer with Denoising Diffusion Probabilistic Models

2022-12-16 · Ziyi Chang, Edmund J. C. Findlay, Haozheng Zhang, Hubert P. H. Shum

Generating realistic motions for digital humans is a core but challenging part of computer animations and games, as human motions are both diverse in content and rich in styles. While the latest deep learning approaches …

DenoisingMotion SynthesisStyle Transfer

Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation

2025-08-23 · Konstantina Nikolaidou, George Retsinas, Giorgos Sfikas, Silvia Cascianelli 외 arxiv

Diffusion-based Handwritten Text Generation (HTG) approaches achieve impressive results on frequent, in-vocabulary words observed at training time and on regular styles. However, they are prone to memorizing training sam…

Text Generation

Diffusion-based Human Motion Style Transfer with Semantic Guidance

2024-03-20 · Lei Hu, Zihao Zhang, Yongjing Ye, Yiwen Xu 외

3D Human motion style transfer is a fundamental problem in computer graphic and animation processing. Existing AdaIN- based methods necessitate datasets with balanced style distribution and content/style labels to train …

Motion Style TransferStyle TransferTransfer Learning

Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion

2026-05-06 · Yiran Qiao, Yiren Lu, Yunlai Zhou, Disheng Liu 외 arxiv

3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects from a single or multiple images. Building on this capability, enabling styl…

Style Transfer3D Generation