paper-with-me

홈 › Papers

PromptCrafter: Crafting Text-to-Image Prompt through Mixed-Initiative Dialogue with LLM

2023-07-18 · Seungho Baek, Hyerin Im, Jiseung Ryu, Juhyeong Park, Takyeon Lee

Text-to-image generation model is able to generate images across a diverse range of subjects and styles based on a single prompt. Recent works have proposed a variety of interaction methods that help users understand the capabilities of models and utilize them. However, how to support users to efficiently explore the model's capability and to create effective prompts are still open-ended research questions. In this paper, we present PromptCrafter, a novel mixed-initiative system that allows step-by-step crafting of text-to-image prompt. Through the iterative process, users can efficiently explore the model's capability, and clarify their intent. PromptCrafter also supports users to refine prompts by answering various responses to clarifying questions generated by a Large Language Model. Lastly, users can revert to a desired step by reviewing the work history. In this workshop paper, we discuss the design process of PromptCrafter and our plans for follow-up studies.

📄 PDF Abstract BibTeX arXiv:2307.08985

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationLanguage ModelingLanguage ModellingLarge Language ModelText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

PromptMap: An Alternative Interaction Style for AI-Based Image Generation

2025-03-12 · Krzysztof Adamkiewicz, Paweł W. Woźniak, Julia Dominiak, Andrzej Romanowski 외

Recent technological advances popularized the use of image generation among the general public. Crafting effective prompts can, however, be difficult for novice users. To tackle this challenge, we developed PromptMap, a …

Image GenerationSemantic SimilaritySemantic Textual Similarity

Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models

2023-04-18 · Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore 외

Text-to-image generative models have demonstrated remarkable capabilities in generating high-quality images based on textual prompts. However, crafting prompts that accurately capture the user's creative intent remains c…

Image GenerationText to Image GenerationText-to-Image Generation

PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement

2024-03-06 · Zhijie Wang, Yuheng Huang, Da Song, Lei Ma 외

The recent advancements in Generative AI have significantly advanced the field of text-to-image generation. The state-of-the-art text-to-image model, Stable Diffusion, is now capable of synthesizing high-quality images w…

Image GenerationImage InpaintingPrompt EngineeringText to Image Generation+1

Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs

2025-10-10 · Yumin Choi, Dongki Kim, Jinheon Baek, Sung Ju Hwang arxiv

Large Language Models (LLMs) have shown remarkable success, and their multimodal expansions (MLLMs) further unlock capabilities spanning images, videos, and other modalities beyond text. However, despite this shift, prom…

SceneCraft: Interactive System for Image Editing via Scene Graph

2026-06-15 · Duc-Manh Phan, Ngoc-Dai Tran, Duy-Khang Do, Tam V. Nguyen 외 arxiv

Recent advances in generative AI have enabled natural language-driven image editing, yet existing systems often fail in complex scenes with multiple interacting objects because they rely heavily on users crafting precise…

Prompt EngineeringImage Editing