paper-with-me

Papers

DiffStyler: Controllable Dual Diffusion for Text-Driven Image Stylization

2022-11-19 · Nisha Huang, Yuxin Zhang, Fan Tang, Chongyang Ma, Haibin Huang, Yong Zhang, WeiMing Dong, Changsheng Xu

Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target style provided by the user. Unlike the previous image-to-image transfer approaches, text-guided stylization progress provides users with a more precise and intuitive way to express the desired style. However, the huge discrepancy between cross-modal inputs/outputs makes it challenging to conduct text-driven image stylization in a typical feed-forward CNN pipeline. In this paper, we present DiffStyler, a dual diffusion processing architecture to control the balance between the content and style of the diffused results. The cross-modal style information can be easily integrated as guidance during the diffusion process step-by-step. Furthermore, we propose a content image-based learnable noise on which the reverse denoising process is based, enabling the stylization results to better preserve the structure information of the content image. We validate the proposed DiffStyler beyond the baseline methods through extensive qualitative and quantitative experiments. Code is available at \url{https://github.com/haha-lisa/Diffstyler}.

📄 PDF Abstract BibTeX arXiv:2211.10682

Code (1)

haha-lisa/Diffstyler 공식 구현 pytorch

Tasks

DenoisingImage StylizationStyle Transfer

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DiffStyler: Diffusion-based Localized Image Style Transfer

2024-03-27 · Shaoxu Li

Image style transfer aims to imbue digital imagery with the distinctive attributes of style targets, such as colors, brushstrokes, shapes, whilst concurrently preserving the semantic integrity of the content. Despite the…

DenoisingStyle Transfer

Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness

2024-01-07 · Sicheng Yang, Zunnan Xu, Haiwei Xue, Yongkang Cheng 외

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation…

Gesture GenerationMotion Generation

HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

2022-11-11 · Kaiduo Zhang, Muyi Sun, Jianxin Sun, Binghao Zhao 외

Text-driven person image generation is an emerging and challenging task in cross-modality image generation. Controllable person image generation promotes a wide range of applications such as digital human interaction and…

Image GenerationRetrievalSentenceVirtual Try-on

3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Models

2023-11-09 · Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao 외

3D content creation via text-driven stylization has played a fundamental challenge to multimedia and graphics community. Recent advances of cross-modal foundation models (e.g., CLIP) have made this problem feasible. Thos…

Image Generation

TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation

2024-12-13 · Xingrui Wang, Xin Li, Yaosi Hu, Hanxin Zhu 외

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the …

Image to Video GenerationObjectVideo Generation