paper-with-me

홈 › Papers

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models

2026-02-26 · Mingde Yao, Zhiyuan You, King-Man Tam, Menglu Wang, Tianfan Xue arxiv

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing the burden of task decomposition and sequencing entirely on the user. To achieve autonomous image editing, we present PhotoAgent, a system that advances image editing through explicit aesthetic planning. Specifically, PhotoAgent formulates autonomous image editing as a long-horizon decision-making problem. It reasons over user aesthetic intent, plans multi-step editing actions via tree search, and iteratively refines results through closed-loop execution with memory and visual feedback, without requiring step-by-step user prompts. To support reliable evaluation in real-world scenarios, we introduce UGC-Edit, an aesthetic evaluation benchmark consisting of 7,000 photos and a learned aesthetic reward model. We also construct a test set containing 1,017 photos to systematically assess autonomous photo editing performance. Extensive experiments demonstrate that PhotoAgent consistently improves both instruction adherence and visual quality compared with baseline methods. The project page is https://mdyao.github.io/PhotoAgent/.

📄 PDF Abstract BibTeX arXiv:2602.22809

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding

2026-03-24 · Lirong Che, Zhenfeng Gan, Yanbo Chen, Junbo Tan 외 arxiv

Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that achieves this by integrating Large Multi…

Spatial Reasoning

Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes

2026-05-28 · Ruixiang Jiang, Chang Wen Chen arxiv

Portrait photography is largely decided before the shutter opens: the subject's pose, the camera configuration, and the lighting devices must be coordinated within the surrounding 3D scene. In contrast, most existing com…

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

2025-12-12 · Chenrui Fan, Yijun Liang, Shweta Bhardwaj, Kwesi Cobbina 외 arxiv

While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks…

Visual Reasoning

Advancing Aesthetic Image Generation via Composition Transfer

2026-05-06 · Kai Zou, Zhiwei Zhao, Bin Liu, Nenghai Yu arxiv

Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, in practice, composition is often coupled with semantics. As a result…

Image Generation

AesRM: Improving Video Aesthetics with Expert-Level Feedback

2026-04-30 · Yujin Han, Yujie Wei, Yefei He, Xinyu Liu 외 arxiv

Despite rapid advances in photorealistic video generation, real-world applications such as filmmaking require video aesthetics, e.g., harmonious colors and cinematic lighting, beyond visual fidelity. Prior work on visual…

Video Generation