paper-with-me

Papers

Prompt-Consistency Image Generation (PCIG): A Unified Framework Integrating LLMs, Knowledge Graphs, and Controllable Diffusion Models

2024-06-24 · Yichen Sun, Zhixuan Chu, Zhan Qin, Kui Ren

The rapid advancement of Text-to-Image(T2I) generative models has enabled the synthesis of high-quality images guided by textual descriptions. Despite this significant progress, these models are often susceptible in generating contents that contradict the input text, which poses a challenge to their reliability and practical deployment. To address this problem, we introduce a novel diffusion-based framework to significantly enhance the alignment of generated images with their corresponding descriptions, addressing the inconsistency between visual output and textual input. Our framework is built upon a comprehensive analysis of inconsistency phenomena, categorizing them based on their manifestation in the image. Leveraging a state-of-the-art large language module, we first extract objects and construct a knowledge graph to predict the locations of these objects in potentially generated images. We then integrate a state-of-the-art controllable image generation model with a visual text generation module to generate an image that is consistent with the original prompt, guided by the predicted object locations. Through extensive experiments on an advanced multimodal hallucination benchmark, we demonstrate the efficacy of our approach in accurately generating the images without the inconsistency with the original prompt. The code can be accessed via https://github.com/TruthAI-Lab/PCIG.

📄 PDF Abstract BibTeX arXiv:2406.16333

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationImage GenerationKnowledge GraphsText Generation

Similar Papers 제목 키워드 기반

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

2026-05-11 · Yu He, Ting Zhu, Yichun Liu, Lichen Ma 외 arxiv

Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remain…

Infinite-Story: A Training-Free Consistent Text-to-Image Generation

2025-11-17 · Jihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo 외 arxiv

We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two …

Text-to-Image GenerationVisual Storytelling

Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models

2025-12-22 · Zixuan Ye, Quande Liu, Cong Wei, Yuanxing Zhang 외 arxiv

Recently, the introduction of Chain-of-Thought (CoT) has largely improved the generation ability of unified models. However, it is observed that the current thinking process during generation mainly focuses on the text c…

ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation

2025-11-14 · Kaishen Wang, Ruibo Chen, Tong Zheng, Heng Huang arxiv

Recent text-to-image (T2I) models have made remarkable progress in generating visually realistic and semantically coherent images. However, they still suffer from randomness and inconsistency with the given prompts, part…

Image Generation

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

2025-12-01 · Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu 외 arxiv

Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to calibrate semantic consistency. However, th…

Image Generation