paper-with-me

홈 › Papers

OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

2025-09-29 · Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang arxiv

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18\% on editing tasks (UniWorld-V1 on ImgEdit-Bench) and 13% on generation tasks (Harmon on GenEval). Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.

📄 PDF Abstract BibTeX arXiv:2509.24900

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationStyle Transfer

Similar Papers 제목 키워드 기반

Data Processing for the OpenGPT-X Model Family

2024-10-11 · Nicolo' Brandizzi, Hammam Abdelwahab, Anirban Bhowmick, Lennard Helmer 외

This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performance multilingual large language models (…

model

Advanced Image Processing for Astronomical Images

2018-12-23 · Diganta Misra, Sparsha Mishra, Bhargav Appasani

Image Processing in Astronomy is a major field of research and involves a lot of techniques pertaining to improve analyzing the properties of the celestial objects or obtaining preliminary inference from the image data. …

Astronomy

GPT-Sentinel: Distinguishing Human and ChatGPT Generated Content

2023-05-13 · Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li 외

This paper presents a novel approach for detecting ChatGPT-generated vs. human-written text using language models. To this end, we first collected and released a pre-processed dataset named OpenGPTText, which consists of…

text-classificationText Classification

Variational Bayesian Framework for Advanced Image Generation with Domain-Related Variables

2023-05-23 · Yuxiao Li, Santiago Mazuelas, Yuan Shen

Deep generative models (DGMs) and their conditional counterparts provide a powerful ability for general-purpose generative modeling of data distributions. However, it remains challenging for existing methods to address a…

Image GenerationImage-to-Image TranslationTranslationUnsupervised Image-To-Image Translation

CongNaMul: A Dataset for Advanced Image Processing of Soybean Sprouts

2023-08-30 · Byunghyun Ban, Donghun Ryu, Su-won Hwang

We present 'CongNaMul', a comprehensive dataset designed for various tasks in soybean sprouts image analysis. The CongNaMul dataset is curated to facilitate tasks such as image classification, semantic segmentation, deco…

image-classificationImage ClassificationSegmentationSemantic Segmentation