paper-with-me

홈 › Papers

GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models

2025-11-14 · Jingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu, Siyuan Li, Linzhuang Sun, Bihui Yu, Conghui He, Lijun Wu, Cheng Tan arxiv

The advent of Unified Multimodal Models (UMMs) signals a paradigm shift in artificial intelligence, moving from passive perception to active, cross-modal generation. Despite their unprecedented ability to synthesize information, a critical gap persists in evaluation: existing benchmarks primarily assess discriminative understanding or unconstrained image generation separately, failing to measure the integrated cognitive process of generative reasoning. To bridge this gap, we propose that geometric construction provides an ideal testbed as it inherently demands a fusion of language comprehension and precise visual generation. We introduce GGBench, a benchmark designed specifically to evaluate geometric generative reasoning. It provides a comprehensive framework for systematically diagnosing a model's ability to not only understand and reason but to actively construct a solution, thereby setting a more rigorous standard for the next generation of intelligent systems. Project website: https://opendatalab-raiser.github.io/GGBench/.

📄 PDF Abstract BibTeX arXiv:2511.11134

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

2026-08-14 · Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi 외 arxiv

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilitie…

multimodal generationSpatial Reasoning3D Reconstruction

Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework

2026-01-10 · Pan Liao, Feng Yang, Di Wu, Jinwen Yu 외 arxiv

Semantic Multi-Object Tracking (SMOT) is evolving from purely geometric localization toward comprehensive video understanding. However, existing paradigms predominantly rely on closed-set interaction tags and fragmented …

Multi-Object Tracking

Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation

2025-09-23 · Yuanhuiyi Lyu, Chi Kit Wong, Chenfei Liao, Lutao Jiang 외 arxiv

Recent works have made notable advancements in enhancing unified models for text-to-image generation through the Chain-of-Thought (CoT). However, these reasoning methods separate the processes of understanding and genera…

Text-to-Image GenerationImage Editing

LeanGeo: Formalizing Competitional Geometry problems in Lean

2025-08-20 · Chendong Song, Zihan Wang, Frederick Pu, Haiming Wang 외 arxiv

Geometry problems are a crucial testbed for AI reasoning capabilities. Most existing geometry solving systems cannot express problems within a unified framework, thus are difficult to integrate with other mathematical fi…

Aether: Geometric-Aware Unified World Modeling

2025-03-24 · Aether Team, Haoyi Zhu, Yifan Wang, Jianjun Zhou 외

The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enab…

Dynamic ReconstructionPredictionSpatial ReasoningTrajectory Planning+3