paper-with-me

홈 › Papers

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

2025-12-01 · Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu, Cheng Tan arxiv

Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to calibrate semantic consistency. However, their reliance on static, single-image generation in training and evaluation leads to overfitting to static pattern matching and semantic fusion, while fundamentally hindering their ability to model dynamic processes that unfold over time. To address these constraints, we propose Envision-a causal event progression benchmark for chained text-to-multi-image generation. Grounded in world knowledge and structured by spatiotemporal causality, it reorganizes existing evaluation dimensions and includes 1,000 four-stage prompts spanning six scientific and humanities domains. To transition evaluation from single images to sequential frames and assess whether models truly internalize world knowledge while adhering to causal-temporal constraints, we introduce Envision-Score, a holistic metric integrating multi-dimensional consistency, physicality, and aesthetics. Comprehensive evaluation of 15 models (10 specialized T2I models, 5 unified models) uncovers: specialized T2I models demonstrate proficiency in aesthetic rendering yet lack intrinsic world knowledge. Unified multimodal models bridge this gap, consistently outperforming specialized counterparts in causal narrative coherence. However, even these unified architectures remain subordinate to closed-source models and struggle to overcome the core challenge of spatiotemporal consistency. This demonstrates that a focus on causally-isolated single images impedes multi-frame reasoning and generation, promoting static pattern matching over dynamic world modeling-ultimately limiting world knowledge internalization, generation.

📄 PDF Abstract BibTeX arXiv:2512.01816

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation

2026-01-21 · Letian Zhang, Sucheng Ren, Yanqing Liu, Xianhang Li 외 arxiv

This paper presents a family of advanced vision encoder, named OpenVision 3, that learns a single, unified visual representation that can serve both image understanding and image generation. Our core architecture is simp…

Contrastive LearningImage Generation

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

2026-05-07 · Zeyu Liu, Zanlin Ni, Yang Yue, Cheng Da 외 arxiv

Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation compo…

Image Generation

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

2026-07-28 · Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li 외 arxiv

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two …

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

2026-06-08 · Lingyi Meng, Zecong Tang, Haoran Li, Tengju Ru 외 arxiv

In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework. Mastering dynamic, multi-turn interleaved image-text dialogues is a crucial task for …

Beyond Correlation: Towards Causal Large Language Model Agents in Biomedicine

2025-05-22 · Adib Bazgir, Amir Habibdoust Lafmajani, Yuwen Zhang

Large Language Models (LLMs) show promise in biomedicine but lack true causal understanding, relying instead on correlations. This paper envisions causal LLM agents that integrate multimodal data (text, images, genomics,…

Causal InferenceDrug DiscoveryLanguage ModelingLanguage Modelling+1