paper-with-me

홈 › Papers

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

2026-05-15 · Xiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang, Yuyao Zhai, Kaitao Lin, Liang Chen, Gai Yuhang, Yuyu Luo, Qiang Wang, Xiaowen Chu arxiv

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows. Existing methods predominantly rely on pixel-based synthesis, which operates in probabilistic pixel spaces and is inherently limited in editability and fidelity. Instead, we propose a new Diagram-as-Code paradigm with symbolic logic that leverages mxGraph Extensible Markup Language (XML) for precise diagram generation and editing. We present VCG-Bench, a unified benchmark for visual-centric \texttt{mxGraph} tasks. VCG-Bench comprises: (1) a taxonomized dataset of 1,449 diverse diagrams spanning 6 domains and 15 sub-domains, (2) a paradigm definition that integrates Generation (Vision-to-Code) and Editability (Code-to-Code), (3) a Tailored Evaluation Protocol employing multi-dimensional metrics such as \texttt{mxGraph} Execution Success Rate, Style Consistency Score (SCS), etc. Experimental results highlight the challenges faced by current State-of-the-Art (SOTA) VLMs in structured fidelity and instruction compliance, reflecting their vision and reasoning capabilities.

📄 PDF Abstract BibTeX arXiv:2605.15677

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

2025-11-11 · Yifei Cao, Yu Liu, Guolong Wang, Zhu Liu 외 arxiv

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages epis…

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

2026-08-24 · Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan 외 arxiv

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects i…

Novel View SynthesisSpatial Reasoning

UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning

2025-09-07 · Huy Le, Nhat Chung, Tung Kieu, Jingkang Yang 외 arxiv

Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typically target either coarse-grained box-…

Video scene graph generationRepresentation Learning

UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation

2025-08-02 · Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka 외 arxiv

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solu…

Motion ForecastingMotion Synthesis

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

2026-04-13 · Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong 외 arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasonin…

Scene UnderstandingSpatial Reasoning