paper-with-me

홈 › Papers

Text-to-Sticker: Style Tailoring Latent Diffusion Models for Human Expression

2023-11-17 · Animesh Sinha, Bo Sun, Anmol Kalia, Arantxa Casanova, Elliot Blanchard, David Yan, Winnie Zhang, Tony Nelli, Jiahui Chen, Hardik Shah, Licheng Yu, Mitesh Kumar Singh, Ankit Ramchandani, Maziar Sanjabi, Sonal Gupta, Amy Bearman, Dhruv Mahajan

We introduce Style Tailoring, a recipe to finetune Latent Diffusion Models (LDMs) in a distinct domain with high visual quality, prompt alignment and scene diversity. We choose sticker image generation as the target domain, as the images significantly differ from photorealistic samples typically generated by large-scale LDMs. We start with a competent text-to-image model, like Emu, and show that relying on prompt engineering with a photorealistic model to generate stickers leads to poor prompt alignment and scene diversity. To overcome these drawbacks, we first finetune Emu on millions of sticker-like images collected using weak supervision to elicit diversity. Next, we curate human-in-the-loop (HITL) Alignment and Style datasets from model generations, and finetune to improve prompt alignment and style alignment respectively. Sequential finetuning on these datasets poses a tradeoff between better style alignment and prompt alignment gains. To address this tradeoff, we propose a novel fine-tuning method called Style Tailoring, which jointly fits the content and style distribution and achieves best tradeoff. Evaluation results show our method improves visual quality by 14%, prompt alignment by 16.2% and scene diversity by 15.3%, compared to prompt engineering the base Emu model for stickers generation.

📄 PDF Abstract BibTeX arXiv:2311.10794

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage GenerationPrompt Engineering

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Animated Stickers: Bringing Stickers to Life with Video Diffusion

2024-02-08 · David Yan, Winnie Zhang, Luxin Zhang, Anmol Kalia 외

We introduce animated stickers, a video diffusion model which generates an animation conditioned on a text prompt and static sticker image. Our model is built on top of the state-of-the-art Emu text-to-image model, with …

SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset

2026-04-29 · Changhyun Roh, Yonghyun Jeong, Jonghyun Lee, Chanho Eom 외 arxiv

Synthesizing a target concept from a single reference image is challenging in diffusion-based personalized text-to-image generation, particularly for sticker personalization where prompts often require explicit attribute…

Text-to-Image GenerationTest-time Adaptation

Sticker820K: Empowering Interactive Retrieval with Stickers

2023-06-12 · Sijie Zhao, Yixiao Ge, Zhongang Qi, Lin Song 외

Stickers have become a ubiquitous part of modern-day communication, conveying complex emotions through visual imagery. To facilitate the development of more powerful algorithms for analyzing stickers, we propose a large-…

Image RetrievalRetrieval

LogoSticker: Inserting Logos into Diffusion Models for Customized Generation

2024-07-18 · Mingkang Zhu, Xi Chen, Zhongdao Wang, Hengshuang Zhao 외

Recent advances in text-to-image model customization have underscored the importance of integrating new concepts with a few examples. Yet, these progresses are largely confined to widely recognized subjects, which can be…

PerSRV: Personalized Sticker Retrieval with Vision-Language Model

2024-10-29 · Heng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma 외

Instant Messaging is a popular means for daily communication, allowing users to send text and stickers. As the saying goes, "a picture is worth a thousand words", so developing an effective sticker retrieval technique is…

Language ModelingLanguage ModellingmodelRetrieval