paper-with-me

Papers

GenSpace: Benchmarking Spatially-Aware Image Generation

2025-05-30 · Zehan Wang, Jiayang Xu, Ziang Zhang, Tianyu Pan, Chao Du, Hengshuang Zhao, Zhou Zhao

Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? We present GenSpace, a novel benchmark and evaluation pipeline to comprehensively assess the spatial awareness of current image generation models. Furthermore, standard evaluations using general Vision-Language Models (VLMs) frequently fail to capture the detailed spatial errors. To handle this challenge, we propose a specialized evaluation pipeline and metric, which reconstructs 3D scene geometry using multiple visual foundation models and provides a more accurate and human-aligned metric of spatial faithfulness. Our findings show that while AI models create visually appealing images and can follow general instructions, they struggle with specific 3D details like object placement, relationships, and measurements. We summarize three core limitations in the spatial perception of current state-of-the-art image generation models: 1) Object Perspective Understanding, 2) Egocentric-Allocentric Transformation and 3) Metric Measurement Adherence, highlighting possible directions for improving spatial intelligence in image generation.

📄 PDF Abstract BibTeX arXiv:2505.24870

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingImage Generation

Similar Papers 제목 키워드 기반

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness

2026-04-29 · Haiyi Qiu, Kaihang Pan, Jiacheng Li, Juncheng Li 외 arxiv

Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image generation. However, these models remain fundamentally limited in sp…

Text-to-Image GenerationImage Editing

EigenGS Representation: From Eigenspace to Gaussian Image Space

2025-03-10 · CVPR 2025 1 · Lo-Wei Tai, Ching-En Li, Cheng-Lin Chen, Chih-Jung Tsai 외

Principal Component Analysis (PCA), a classical dimensionality reduction technique, and 2D Gaussian representation, an adaptation of 3D Gaussian Splatting for image representation, offer distinct approaches to modeling v…

Dimensionality Reduction

Learning Laplacian Eigenspace with Mass-Aware Neural Operators on Point Clouds

2026-05-23 · Zherui Yang, Tao Du, Ligang Liu arxiv

The eigendecomposition of the Laplace--Beltrami Operator (LBO) is fundamental to geometric analysis, yet computing its low-frequency eigenmodes remains a significant bottleneck due to the high cost of iterative solvers o…

Zero-shot GeneralizationPoint Clouds

Automatic Spatially-aware Fashion Concept Discovery

2017-08-03 · ICCV 2017 10 · Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang 외

This paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites. We first fine-tune GoogleNet by jointly modeling clothing images and their correspo…

AttributeClusteringImage Retrieval with Multi-Modal QueryRetrieval

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

2026-06-29 · Filippo Ruffini, Marco Salmé, Rosa Sicilia, Valerio Guarrasi 외 arxiv

Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do no…