paper-with-me

홈 › Papers

From "What" to "How": Constrained Reasoning for Autoregressive Image Generation

2026-03-03 · Ruxue Yan, Xubo Liu, Wenya Guo, Zhengkun Zhang, Ying Zhang, Xiaojie Yuan arxiv

Autoregressive image generation has seen recent improvements with the introduction of chain-of-thought and reinforcement learning. However, current methods merely specify "What" details to depict by rewriting the input prompt, yet fundamentally fail to reason about "How" to structure the overall image. This inherent limitation gives rise to persistent issues, such as spatial ambiguity directly causing unrealistic object overlaps. To bridge this gap, we propose CoR-Painter, a novel framework that pioneers a "How-to-What" paradigm by introducing Constrained Reasoning to guide the autoregressive generation. Specifically, it first deduces "How to draw" by deriving a set of visual constraints from the input prompt, which explicitly govern spatial relationships, key attributes, and compositional rules. These constraints steer the subsequent generation of a detailed description "What to draw", providing a structurally sound and coherent basis for accurate visual synthesis. Additionally, we introduce a Dual-Objective GRPO strategy that specifically optimizes the textual constrained reasoning and visual projection processes to ensure the coherence and quality of the entire generation pipeline. Extensive experiments on T2I-CompBench, GenEval, and WISE demonstrate that our method achieves state-of-the-art performance, with significant improvements in spatial metrics (e.g., +5.41% on T2I-CompBench).

📄 PDF Abstract BibTeX arXiv:2603.02712

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Generation

Similar Papers 제목 키워드 기반

IRIS: Intrinsic Reward Image Synthesis

2025-09-29 · Yihang Chen, Yuanhao Ban, Yunqi Hong, Cho-Jui Hsieh arxiv

Despite the success of Reinforcement Learning from Human Feedback (RLHF) in language reasoning, its application to autoregressive Text-to-Image (T2I) generation is often constrained by the limited availability of human p…

Reinforcement LearningImage Generation

Autoregressive Image Generation Guided by Chains of Thought

2025-02-24 · Miaomiao Cai, Guanjie Wang, Wei Li, Zhijun Tu 외

In the field of autoregressive (AR) image generation, models based on the 'next-token prediction' paradigm of LLMs have shown comparable performance to diffusion models by reducing inductive biases. However, directly app…

Image GenerationLogical Reasoning

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers

2026-06-27 · Achin Jain, Jie An, Siddharth Chaudhary, Davide Modolo arxiv

Leveraging capabilities of large language models (LLMs) in text-to-image (T2I) synthesis is an important research direction. In this work we investigate whether the knowledge of a frozen LLM can be effectively utilized i…

Image Generation

Let's Verify and Reinforce Image Generation Step by Step

2025-01-01 · CVPR 2025 1 · Renrui Zhang, Chengzhuo Tong, Zhizheng Zhao, Ziyu Guo 외

Chain-of-Thought (CoT) reasoning has been extensively explored in large models to tackle complex understanding tasks. However, it still remains an open question whether such strategies can be applied to verifying and…

Image Generation

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

2025-01-23 · Ziyu Guo, Renrui Zhang, Chengzhuo Tong, Zhizheng Zhao 외

Chain-of-Thought (CoT) reasoning has been extensively explored in large models to tackle complex understanding tasks. However, it still remains an open question whether such strategies can be applied to verifying and rei…

Image GenerationText-to-Image Generation