paper-with-me

홈 › Papers

ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation

2024-10-02 · Rinon Gal, Adi Haviv, Yuval Alaluf, Amit H. Bermano, Daniel Cohen-Or, Gal Chechik

The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting effective workflows requires significant expertise, owing to the large number of available components, their complex inter-dependence, and their dependence on the generation prompt. Here, we introduce the novel task of prompt-adaptive workflow generation, where the goal is to automatically tailor a workflow to each user prompt. We propose two LLM-based approaches to tackle this task: a tuning-based method that learns from user-preference data, and a training-free method that uses the LLM to select existing flows. Both approaches lead to improved image quality when compared to monolithic models or generic, prompt-independent workflows. Our work shows that prompt-dependent flow prediction offers a new pathway to improving text-to-image generation quality, complementing existing research directions in the field.

📄 PDF Abstract BibTeX arXiv:2410.01731

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

3DALL-E: Integrating Text-to-Image AI in 3D Design Workflows

2022-10-20 · Vivian Liu, Jo Vermeulen, George Fitzmaurice, Justin Matejka

Text-to-image AI are capable of generating novel images for inspiration, but their applications for 3D design workflows and how designers can build 3D models using AI-provided inspiration have not yet been explored. To i…

Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing

2025-11-15 · Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao, Pascal Poupart 외 arxiv

Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative workflows. We introduce Image-POSER, a r…

Text-to-Image GenerationReinforcement Learning

EvoFlow: Evolving Diverse Agentic Workflows On The Fly

2025-02-11 · Guibin Zhang, Kaijie Chen, Guancheng Wan, Heng Chang 외

The past two years have witnessed the evolution of large language model (LLM)-based multi-agent systems from labor-intensive manual design to partial automation (\textit{e.g.}, prompt engineering, communication topology)…

Large Language ModelPrompt EngineeringTAG

Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning

2025-02-12 · Zijian He, Reyna Abhyankar, Vikranth Srivatsa, Yiying Zhang

Today's gen-AI workflows that involve multiple ML model calls, tool/API calls, data retrieval, or generic code execution are often tuned manually in an ad-hoc way that is both time-consuming and error-prone. In this pape…

RAGText to SQLText-To-SQL

Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation

2025-12-20 · Ziyu Zhang, Yi Yu, Simeng Zhu, Ahmed Aly 외 arxiv

Accurate segmentation of anatomical structures in medical images is essential for diagnosis and treatment planning. While recent interactive segmentation foundation models enhance generalization through large-scale multi…

Medical Image SegmentationInteractive Segmentation