paper-with-me

Papers

UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

2026-02-02 · Dianyi Wang, Chaofan Ma, Feng Han, Size Wu, Wei Song, Yibin Wang, Zhixiong Zhang, Tianhang Wang, Siyuan Wang, Zhongyu Wei, Jiaqi Wang arxiv

Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoning steps. To address this, we propose UniReason, a unified framework that harmonizes these two tasks through two complementary reasoning paradigms. We incorporate world knowledge-enhanced textual reasoning into generation to infer implicit knowledge, and leverage editing capabilities for fine-grained editing-like visual refinement to further correct visual errors via self-reflection. This approach unifies generation and editing within a shared architecture, mirroring the human cognitive process of planning followed by refinement. We support this framework by systematically constructing a large-scale reasoning-centric dataset (~300k samples) covering five major knowledge domains (e.g., cultural commonsense, physics, etc.) for textual reasoning, alongside an agent-generated corpus for visual refinement. Extensive experiments demonstrate that UniReason achieves advanced performance on reasoning-intensive benchmarks such as WISE, KrisBench and UniREditBench, while maintaining superior general synthesis capabilities.

📄 PDF Abstract BibTeX arXiv:2602.02437

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationImage Editing

Similar Papers 제목 키워드 기반

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

2026-06-10 · Mengzhuo Chen, Yan Shu, Chi Liu, Hongming Piao 외 arxiv

We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a common reasoning interface. We introduce UniReason-Med, a single-chec…

Reinforcement Learning

Large Language Models are Universal Reasoners for Visual Generation

2026-05-05 · Sucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song 외 arxiv

Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM backbone handles both visual understanding and generation. Despite the …

Text-to-Image Generation

Pandora: Leveraging Code-driven Knowledge Transfer for Unified Structured Knowledge Reasoning

2025-08-25 · Yongrui Chen, Junhao He, Linbo Fu, Shenyu Zhang 외 arxiv

Unified Structured Knowledge Reasoning (USKR) aims to answer natural language questions by using structured sources such as tables, databases, and knowledge graphs in a unified way. Existing USKR methods rely on task-spe…

Knowledge Graphs

Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks

2026-02-02 · Bohan Zeng, Kaixin Zhu, Daili Hua, Bozhou Li 외 arxiv

World models have emerged as a critical frontier in AI research, aiming to enhance large models by infusing them with physical dynamics and world knowledge. The core objective is to enable agents to understand, predict, …

AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models

2026-01-02 · Jintao Lin, Bowen Dong, Weikang Shi, Chenyang Lei 외 arxiv

The capability of Unified Multimodal Models (UMMs) to apply world knowledge across diverse tasks remains a critical, unresolved challenge. Existing benchmarks fall short, offering only siloed, single-task evaluations wit…