paper-with-me

홈 › Papers

Compositional Image Synthesis with Inference-Time Scaling

2025-10-28 · Minsuk Ji, Sanghyeok Lee, Namhyuk Ahn arxiv

Despite their impressive realism, modern text-to-image models still struggle with compositionality, often failing to render accurate object counts, attributes, and spatial relations. To address this challenge, we present a training-free framework that combines an object-centric approach with self-refinement to improve layout faithfulness while preserving aesthetic quality. Specifically, we leverage large language models (LLMs) to synthesize explicit layouts from input prompts, and we inject these layouts into the image generation process, where a object-centric vision-language model (VLM) judge reranks multiple candidates to select the most prompt-aligned outcome iteratively. By unifying explicit layout-grounding with self-refine-based inference-time scaling, our framework achieves stronger scene alignment with prompts compared to recent text-to-image models. The code are available at https://github.com/gcl-inha/ReFocus.

📄 PDF Abstract BibTeX arXiv:2510.24133

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

2025-11-23 · Shijian Wang, Runhao Fu, Siyi Zhao, Qingqin Zhan 외 arxiv

Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scenes containing multiple objects that exhi…

Text-to-Image GenerationReinforcement Learning

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing

2025-07-16 · Shreya Kadambi, Risheek Garrepalli, Shubhankar Borse, Munawar Hyatt 외 arxiv

Despite the remarkable success of diffusion models in text-to-image generation, their effectiveness in grounded visual editing and compositional control remains challenging. Motivated by advances in self-supervised learn…

Self-Supervised LearningText-to-Image Generation

All-in-One Conditioning for Text-to-Image Synthesis

2026-02-09 · Hirunima Jayasekara, Chuong Huynh, Yixuan Ren, Christabel Acquaye 외 arxiv

Accurate interpretation and visual representation of complex prompts involving multiple objects, attributes, and spatial relationships is a critical challenge in text-to-image synthesis. Despite recent advancements in ge…

IV-Mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video Synthesis

2024-10-05 · Shitong Shao, Zikai Zhou, Lichen Bai, Haoyi Xiong 외

The multi-step sampling mechanism, a key feature of visual diffusion models, has significant potential to replicate the success of OpenAI's Strawberry in enhancing performance by increasing the inference computational co…

Text-to-Video Generation

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

2025-02-06 · Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang 외

Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and inference-time compute. However, current …

Speech Synthesis