paper-with-me

Papers

AesFormer: Transform Everyday Photos into Beautiful Memories

2026-05-21 · Tianxiang Du, Hulingxiao He, Yuxin Peng arxiv

In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that existing retouching and portrait enhancement methods cannot fix. We formulate Aesthetic Photo Reconstruction (APR) as improving a photo's aesthetic quality via structural reconstruction while preserving subject identity and scene semantics. Although recent advances in image editing models make APR feasible, they often lack aesthetic understanding, yielding edits that are semantically plausible yet aesthetically weak. To address this, we propose AesFormer, a two-stage framework that decouples aesthetic planning from image editing. In Stage 1, an aesthetic action model (AesThinker) analyzes the input along seven progressive photographic dimensions and outputs executable editing actions; we further apply GRPO-A to encourage broad exploration over diverse action plans beyond SFT. In Stage 2, an action-conditioned editor (AesEditor) performs structural edits guided by these actions. To support APR, we build a video-based corpus-mining pipeline (VCMP) and construct AesRecon, a benchmark of 9,071 strictly aligned (poor, good) image pairs. Experiments show that AesFormer substantially improves APR performance and is competitive with Nano Banana Pro. Code is available at https://github.com/PKU-ICST-MIPL/AesFormer_ICML2026.

📄 PDF Abstract BibTeX arXiv:2605.22126

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Beautiful and damned. Combined effect of content quality and social ties on user engagement

2017-11-01 · Luca M. Aiello, Rossano Schifanella, Miriam Redi, Stacey Svetlichnaya 외

User participation in online communities is driven by the intertwinement of the social network structure with the crowd-generated content that flows along its links. These aspects are rarely explored jointly and at scale…

Recommendation Systems

PhotoFramer: Multi-modal Image Composition Instruction

2025-11-30 · Zhiyuan You, Ke Wang, He Zhang, Xin Cai 외 arxiv

Composition matters during the photo-taking process, yet many casual users struggle to frame well-composed images. To provide composition guidance, we introduce PhotoFramer, a multi-modal composition instruction framewor…

Where to Buy It: Matching Street Clothing Photos in Online Shops

2015-12-01 · ICCV 2015 12 · M. Hadi Kiapour, Xufeng Han, Svetlana Lazebnik, Alexander C. Berg 외

In this paper, we define a new task, Exact Street to Shop, where our goal is to match a real-world example of a garment item to the same item in an online shop. This is an extremely challenging task due to visual differe…

Deep LearningRetrieval

Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps

2025-12-19 · Sandeep Mishra, Yasamin Jafarian, Andreas Lugmayr, Yingwei Li 외 arxiv

Photographs of people taken by professional photographers typically present the person in beautiful lighting, with an interesting pose, and flattering quality. This is unlike common photos people take of themselves in un…

Virtual Try-on

Person Recognition in Personal Photo Collections

2015-09-11 · ICCV 2015 12 · Seong Joon Oh, Rodrigo Benenson, Mario Fritz, Bernt Schiele

Recognising persons in everyday photos presents major challenges (occluded faces, different clothing, locations, etc.) for machine vision. We propose a convnet based person recognition system on which we provide an in-de…

InformativenessPerson Recognition