paper-with-me

홈 › Papers

Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models

2023-04-18 · Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, Tovi Grossman

Text-to-image generative models have demonstrated remarkable capabilities in generating high-quality images based on textual prompts. However, crafting prompts that accurately capture the user's creative intent remains challenging. It often involves laborious trial-and-error procedures to ensure that the model interprets the prompts in alignment with the user's intention. To address the challenges, we present Promptify, an interactive system that supports prompt exploration and refinement for text-to-image generative models. Promptify utilizes a suggestion engine powered by large language models to help users quickly explore and craft diverse prompts. Our interface allows users to organize the generated images flexibly, and based on their preferences, Promptify suggests potential changes to the original prompt. This feedback loop enables users to iteratively refine their prompts and enhance desired features while avoiding unwanted ones. Our user study shows that Promptify effectively facilitates the text-to-image workflow and outperforms an existing baseline tool widely used for text-to-image generation.

📄 PDF Abstract BibTeX arXiv:2304.09337

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions

2024-02-05 · Yiyuan Zhang, Yuhao Kang, Zhixin Zhang, Xiaohan Ding 외

We introduce $\textit{InteractiveVideo}$, a user-centric framework for video generation. Different from traditional generative approaches that operate based on user-provided images or text, our framework is designed for …

Video Generation

RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance

2025-10-26 · Jiuniu Wang, Gongjie Zhang, Quanhao Qian, Junlong Gao 외 arxiv

Scalable Vector Graphics (SVGs) are fundamental to digital design and robot control, encoding not only visual structure but also motion paths in interactive drawings. In this work, we introduce RoboSVG, a unified multimo…

CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation

2025-08-25 · Mingyue Yang, Dianxi Shi, Jialu Zhou, Xinyu Wei 외 arxiv

In Text-to-Image (T2I) generation, the complexity of entities and their intricate interactions pose a significant challenge for T2I method based on diffusion model: how to effectively control entity and their interaction…

Text-to-Image Generation

Interactive Text Generation

2023-03-02 · Felix Faltings, Michel Galley, Baolin Peng, Kianté Brantley 외

Users interact with text, image, code, or other editors on a daily basis. However, machine learning models are rarely trained in the settings that reflect the interactivity between users and their editor. This is underst…

Image GenerationImitation LearningText Generation

CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

2023-11-30 · Zineng Tang, ZiYi Yang, Mahmoud Khademi, Yang Liu 외

We present CoDi-2, a versatile and interactive Multimodal Large Language Model (MLLM) that can follow complex multimodal interleaved instructions, conduct in-context learning (ICL), reason, chat, edit, etc., in an any-to…

Image GenerationIn-Context LearningLanguage ModelingLanguage Modelling+3