paper-with-me

Papers

Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation

2023-05-18 · Wanrong Zhu, Xinyi Wang, Yujie Lu, Tsu-Jui Fu, Xin Eric Wang, Miguel Eckstein, William Yang Wang

The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users. Despite the advancements of T2I models, a common issue encountered by users is the need for repetitive editing of input prompts in order to receive a satisfactory image, which is time-consuming and labor-intensive. Given the demonstrated text generation power of large-scale language models, such as GPT-k, we investigate the potential of utilizing such models to improve the prompt editing process for T2I generation. We conduct a series of experiments to compare the common edits made by humans and GPT-k, evaluate the performance of GPT-k in prompting T2I, and examine factors that may influence this process. We found that GPT-k models focus more on inserting modifiers while humans tend to replace words and phrases, which includes changes to the subject matter. Experimental results show that GPT-k are more effective in adjusting modifiers rather than predicting spontaneous changes in the primary subject matters. Adopting the edit suggested by GPT-k models may reduce the percentage of remaining edits by 20-30%.

📄 PDF Abstract BibTeX arXiv:2305.11317

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization

2025-11-06 · Connor Dunlop, Matthew Zheng, Kavana Venkatesh, Pinar Yanardag arxiv

Text-to-image (T2I) diffusion models have made remarkable strides in generating and editing high-fidelity images from text. Yet, these models remain fundamentally generic, failing to adapt to the nuanced aesthetic prefer…

Graph Neural NetworkImage Editing

Collaborative Score Distillation for Consistent Visual Editing

2023-09-21 · NeurIPS 2023 11

Generative priors of large-scale text-to-image diffusion models enable a wide range of new generation and editing applications on diverse visual modalities. However, when adapting these priors to complex visual modalitie…

Collaborative Score Distillation for Consistent Visual Synthesis

2023-07-04 · Subin Kim, Kyungmin Lee, June Suk Choi, Jongheon Jeong 외

Generative priors of large-scale text-to-image diffusion models enable a wide range of new generation and editing applications on diverse visual modalities. However, when adapting these priors to complex visual modalitie…

CCA: Collaborative Competitive Agents for Image Editing

2024-01-23 · Tiankai Hang, Shuyang Gu, Dong Chen, Xin Geng 외

This paper presents a novel generative model, Collaborative Competitive Agents (CCA), which leverages the capabilities of multiple Large Language Models (LLMs) based agents to execute complex tasks. Drawing inspiration f…

Collaborative Diffusion for Multi-Modal Face Generation and Editing

2023-04-20 · CVPR 2023 1 · Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei Liu

Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condit…

DenoisingFace Generation