paper-with-me

홈 › Papers

ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting

2024-11-26 · CVPR 2025 1 · Chengyou Jia, Changliang Xia, Zhuohang Dang, Weijia Wu, Hangwei Qian, Minnan Luo

Despite the significant advancements in text-to-image (T2I) generative models, users often face a trial-and-error challenge in practical scenarios. This challenge arises from the complexity and uncertainty of tedious steps such as crafting suitable prompts, selecting appropriate models, and configuring specific arguments, making users resort to labor-intensive attempts for desired images. This paper proposes Automatic T2I generation, which aims to automate these tedious steps, allowing users to simply describe their needs in a freestyle chatting way. To systematically study this problem, we first introduce ChatGenBench, a novel benchmark designed for Automatic T2I. It features high-quality paired data with diverse freestyle inputs, enabling comprehensive evaluation of automatic T2I models across all steps. Additionally, recognizing Automatic T2I as a complex multi-step reasoning task, we propose ChatGen-Evo, a multi-stage evolution strategy that progressively equips models with essential automation skills. Through extensive evaluation across step-wise accuracy and image quality, ChatGen-Evo significantly enhances performance over various baselines. Our evaluation also uncovers valuable insights for advancing automatic T2I. All our data, code, and models will be available in \url{https://chengyou-jia.github.io/ChatGen-Home}

📄 PDF Abstract BibTeX arXiv:2411.17176

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Interactive Data Synthesis for Systematic Vision Adaptation via LLMs-AIGCs Collaboration

2023-05-22 · Qifan Yu, Juncheng Li, Wentao Ye, Siliang Tang 외

Recent text-to-image generation models have shown promising results in generating high-fidelity photo-realistic images. In parallel, the problem of data scarcity has brought a growing interest in employing AIGC technolog…

Data AugmentationImage GenerationPrompt EngineeringText to Image Generation+1

Freestyle Layout-to-Image Synthesis

2023-03-25 · CVPR 2023 1 · Han Xue, Zhiwu Huang, Qianru Sun, Li Song 외

Typical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle capability of the model, i.e., how far c…

image-classificationImage ClassificationImage GenerationLayout-to-Image Generation+2

PERSONACHATGEN: Generating Personalized Dialogues using GPT-3

2022-10-01 · CCGPK (COLING) 2022 10 · Young-Jun Lee, Chae-Gyun Lim, Yunsu Choi, Ji-Hui Lm 외

Recently, many prior works have made their own agents generate more personalized and engaging responses using personachat. However, since this dataset is frozen in 2018, the dialogue agents trained on this dataset would …

Sentence

FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models

2024-01-28 · Feihong He, Gang Li, Fuhui Sun, Mengyuan Zhang 외

The rapid development of generative diffusion models has significantly advanced the field of style transfer. However, most current style transfer methods based on diffusion models typically involve a slow iterative optim…

DecoderStyle Transfer

Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation

2024-08-28 · Ziqian Ning, Shuai Wang, Yuepeng Jiang, Jixun Yao 외

Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits…

Language ModelingLanguage Modelling