paper-with-me

Papers

Understanding Generative AI Capabilities in Everyday Image Editing Tasks

2025-05-22 · Mohammad Reza Taesiri, Brandon Collins, Logan Bolton, Viet Dac Lai, Franck Dernoncourt, Trung Bui, Anh Totti Nguyen

Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing actions do they want to perform (e.g., removing or stylizing the subject)? Do people prefer precise edits with predictable outcomes or highly creative ones? By understanding the characteristics of real-world requests and the corresponding edits made by freelance photo-editing wizards, can we draw lessons for improving AI-based editors and determine which types of requests can currently be handled successfully by AI editors? In this paper, we present a unique study addressing these questions by analyzing 83k requests from the past 12 years (2013-2025) on the Reddit community, which collected 305k PSR-wizard edits. According to human ratings, approximately only 33% of requests can be fulfilled by the best AI editors (including GPT-4o, Gemini-2.0-Flash, SeedEdit). Interestingly, AI editors perform worse on low-creativity requests that require precise editing than on more open-ended tasks. They often struggle to preserve the identity of people and animals, and frequently make non-requested touch-ups. On the other side of the table, VLM judges (e.g., o1) perform differently from human judges and may prefer AI edits more than human edits. Code and qualitative examples are available at: https://psrdataset.github.io

📄 PDF Abstract BibTeX arXiv:2505.16181

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation

2025-09-23 · Yuanhuiyi Lyu, Chi Kit Wong, Chenfei Liao, Lutao Jiang 외 arxiv

Recent works have made notable advancements in enhancing unified models for text-to-image generation through the Chain-of-Thought (CoT). However, these reasoning methods separate the processes of understanding and genera…

Text-to-Image GenerationImage Editing

Generative Visual Instruction Tuning

2024-06-17 · Jefferson Hernandez, Ruben Villegas, Vicente Ordonez

We propose to use automatically generated instruction-following data to improve the zero-shot capabilities of a large multimodal model with additional support for generative and image editing tasks. We achieve this by cu…

Image GenerationImage-text matchingInstruction FollowingLanguage Modeling+4

DanceOPD: On-Policy Generative Field Distillation

2026-06-25 · Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong 외 arxiv

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often confl…

Image Generation

AesFormer: Transform Everyday Photos into Beautiful Memories

2026-05-21 · Tianxiang Du, Hulingxiao He, Yuxin Peng arxiv

In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that existing retouching and portrait enhancement methods cannot fix. We fo…

Image Editing

OmniCreator: Self-Supervised Unified Generation with Universal Editing

2024-12-03 · Haodong Chen, Lan Wang, Harry Yang, Ser-Nam Lim

We introduce OmniCreator, a novel framework that can conduct text-prompted unified (image+video) generation as well as editing all in one place. OmniCreator acquires generative and universal editing capabilities in a sel…

DenoisingSemantic correspondenceVideo EditingVideo Generation