paper-with-me

Papers

Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation

2026-03-18 · Jiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang, Haibo Qiu, Lihuo He, Dengpan Ye, Xinbo Gao, Jing Zhang arxiv

We present Omni-I2C, a comprehensive benchmark designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executable code. We argue that this task represents a non-trivial challenge for the current generation of LMMs: it demands an unprecedented synergy between high-fidelity visual perception -- to parse intricate spatial hierarchies and symbolic details -- and precise generative expression -- to synthesize syntactically sound and logically consistent code. Unlike traditional descriptive tasks, Omni-I2C requires a holistic understanding where any minor perceptual hallucination or coding error leads to a complete failure in visual reconstruction. Omni-I2C features 1080 meticulously curated samples, defined by its breadth across subjects, image modalities, and programming languages. By incorporating authentic user-sourced cases, the benchmark spans a vast spectrum of digital content -- from scientific visualizations to complex symbolic notations -- each paired with executable reference code. To complement this diversity, our evaluation framework provides necessary depth; by decoupling performance into perceptual fidelity and symbolic precision, it transcends surface-level accuracy to expose the granular structural failures and reasoning bottlenecks of current LMMs. Our evaluation reveals a substantial performance gap among leading LMMs; even state-of-the-art models struggle to preserve structural integrity in complex scenarios, underscoring that multimodal code generation remains a formidable challenge. Data and code are available at https://github.com/MiliLab/Omni-I2C.

📄 PDF Abstract BibTeX arXiv:2603.17508

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

2025-12-11 · Tsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang 외 arxiv

Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen contexts. However, existing methods rely on holistic embeddings from general-…

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

2025-08-25 · Jianwen Tan, Huiyao Zhang, Rui Xiong, Han Zhou 외 arxiv

Camouflaged Object Segmentation (COS) poses a significant challenge due to the intrinsic high similarity between targets and backgrounds, demanding models capable of profound holistic understanding beyond superficial cue…

Camouflaged Object SegmentationMedical Image SegmentationScene Understanding

OmniPerson: Unified Identity-Preserving Pedestrian Generation

2025-12-02 · Changxiao Ma, Chao Yuan, Xincheng Shi, Yuzhuo Ma 외 arxiv

Person re-identification (ReID) suffers from a lack of large-scale high-quality training data due to challenges in data privacy and annotation costs. While previous approaches have explored pedestrian generation for data…

Person Re-IdentificationImage Super-ResolutionData AugmentationVideo Generation

PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

2026-05-19 · Xin Zhang, Yabo Chen, Yijie Fang, Wanying Qu 외 arxiv

Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing video generation methods provide minimal physical control, and single-…

3D ReconstructionScene GenerationVideo Generation3D Generation

Kling-Omni Technical Report

2025-12-18 · Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du 외 arxiv

We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional …

Instruction FollowingVideo Generation