paper-with-me

홈 › Papers

CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation

2025-12-03 · Ruoxuan Zhang, Bin Wen, Hongxia Xie, Yi Yao, Songhan Zuo, Jian-Yu Jiang-Lin, Hong-Han Shuai, Wen-Huang Cheng arxiv

Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusion models have shown strong capabilities in text-to-image generation, they struggle to handle structured multi-step scenarios like recipe illustration. Additionally, current recipe illustration methods are unable to adjust to the natural variability in recipe length, generating a fixed number of images regardless of the actual instructions structure. To address these limitations, we present CookAnything, a flexible and consistent diffusion-based framework that generates coherent, semantically distinct image sequences from textual cooking instructions of arbitrary length. The framework introduces three key components: (1) Step-wise Regional Control (SRC), which aligns textual steps with corresponding image regions within a single denoising process; (2) Flexible RoPE, a step-aware positional encoding mechanism that enhances both temporal coherence and spatial diversity; and (3) Cross-Step Consistency Control (CSCC), which maintains fine-grained ingredient consistency across steps. Experimental results on recipe illustration benchmarks show that CookAnything performs better than existing methods in training-based and training-free settings. The proposed framework supports scalable, high-quality visual synthesis of complex multi-step instructions and holds significant potential for broad applications in instructional media, and procedural content creation.

📄 PDF Abstract BibTeX arXiv:2512.03540

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

ADOPT: Adaptive Dependency-Guided Joint Prompt Optimization for Multi-Step LLM Pipelines

2025-12-31 · Minjun Zhao, Xinyu Zhang, Shuai Zhang, Deyang Li 외 arxiv

Multi-step LLM pipelines can solve complex tasks, but jointly optimizing prompts across steps remains challenging due to missing step-level supervision and inter-step dependency. We propose ADOPT, an adaptive dependency-…

Detecting Multiple Structural Breaks in Systems of Linear Regression Equations with Integrated and Stationary Regressors

2022-01-14 · Karsten Schweikert

In this paper, we propose a two-step procedure based on the group LASSO estimator in combination with a backward elimination algorithm to detect multiple structural breaks in linear regressions with multivariate response…

regression

RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation

2026-05-11 · Zhen Zhang, Wanjing Zhou, Juncheng Li, Hao Fei 외 arxiv

Compared with individual agents, large language model based multi-agent systems have shown great capabilities consistently across diverse tasks, including code generation, mathematical reasoning, and planning, etc. Despi…

Mathematical ReasoningCode Generation

Unifying Light Field Perception with Field of Parallax

2025-03-02 · Fei Teng, Buyin Deng, Boyuan Zheng, Kai Luo 외

Field of Parallax (FoP)}, a spatial field that distills the common features from different LF representations to provide flexible and consistent support for multi-task learning. FoP is built upon three core features--pro…

Multi-Task Learningobject-detectionObject DetectionSalient Object Detection+1

Optimization-based modelling and game-theoretic framework for techno-economic analysis of demand-side flexibility: a real case study

2021-10-14 · Timur Sayfutdinov, Charalampos Patsios, David Greenwood, Meltem Peker 외

This paper proposes a two-step framework for techno-economic analysis of a demand-side flexibility service in distribution networks. Step one applies optimization-based modelling to propose a generic problem formulation …