paper-with-me

Papers

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

2026-03-26 · Yan Li, Zezi Zeng, Ziwei Zhou, Xin Gao, Muzhao Tian, Yifan Yang, Mingxi Cheng, Qi Dai, Yuqing Yang, Lili Qiu, Zhendong Wang, Zhengyuan Yang, Xue Yang, Lijuan Wang, Ji Li, Chong Luo arxiv

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly focus on natural image synthesis and fail to systematically evaluate models under the structured and multi-constraint requirements of real-world commercial design tasks. In this work, we introduce BizGenEval, a systematic benchmark for commercial visual content generation. The benchmark spans five representative document types: slides, charts, webpages, posters, and scientific figures, and evaluates four key capability dimensions: text rendering, layout control, attribute binding, and knowledge-based reasoning, forming 20 diverse evaluation tasks. BizGenEval contains 400 carefully curated prompts and 8000 human-verified checklist questions to rigorously assess whether generated images satisfy complex visual and semantic constraints. We conduct large-scale benchmarking on 26 popular image generation systems, including state-of-the-art commercial APIs and leading open-source models. The results reveal substantial capability gaps between current generative models and the requirements of professional visual content creation. We hope BizGenEval serves as a standardized benchmark for real-world commercial visual content generation.

📄 PDF Abstract BibTeX arXiv:2603.25732

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Ad-Net: Audio-Visual Convolutional Neural Network for Advertisement Detection In Videos

2018-06-22 · Shervin Minaee, Imed Bouazizi, Prakash Kolan, Hossein Najafzadeh

Personalized advertisement is a crucial task for many of the online businesses and video broadcasters. Many of today's broadcasters use the same commercial for all customers, but as one can imagine different viewers have…

ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services

2026-04-27 · Fengxian Ji, Jingpu Yang, Zirui Song, Lang Gao 외 arxiv

Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their performance on paid, real-world design projects remains uncertain. We…

Image Generation

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

2026-04-07 · Yanming Xiu, Zhengyuan Jiang, Neil Zhenqiang Gong, Maria Gorlatova arxiv

Augmented reality (AR) has rapidly expanded over the past decade. As AR becomes increasingly integrated into daily life, its security and reliability emerge as critical challenges. Among various threats, contradictory vi…

A Multimodal Analysis of Influencer Content on Twitter

2023-09-06 · Danae Sánchez Villegas, Catalina Goanta, Nikolaos Aletras

Influencer marketing involves a wide range of strategies in which brands collaborate with popular content creators (i.e., influencers) to leverage their reach, trust, and impact on their audience to promote and endorse p…

Marketing

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

2026-08-28 · Ruijie Jian, Benlei Cui, Ting Ma, Haidong Ding 외 arxiv

Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content platforms where users continuously revise t…