paper-with-me

홈 › Papers

Evolving to the Aesthetics of a Vision-Language Model

2026-05-27 · Stephen James Krol, Jon McCormack arxiv

Evolutionary systems have demonstrated remarkable results in creative domains, with recent applications in generative typography, design, and music. However, an open problem remains in designing fitness functions that effectively capture the desired aesthetics of abstract outputs. In this work, we explore two methods for evaluating the aesthetics of a population using Vision-Language Models (VLMs). The first method uses CLIP-IQA to predict an aesthetic score for each design. The second method instead pits candidates against each other, with winners determined by a VLM using a custom prompt specified by the user. The outcomes of these pairwise comparisons are then used to estimate a population ranking via the Glicko rating system. We present these methods in the context of a case study using a custom generative system and compare the resulting rankings with an artist's aesthetic ranking and those produced by other aesthetic evaluation techniques. Additionally, we document the artist's experience using these approaches to evolve designs, critically analysing the strengths and weaknesses of both methods.

📄 PDF Abstract BibTeX arXiv:2606.00112

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Image Aesthetics Assessment via Learnable Queries

2023-09-06 · Zhiwei Xiong, Yunfan Zhang, Zhiqi Shen, Peiran Ren 외

Image aesthetics assessment (IAA) aims to estimate the aesthetics of images. Depending on the content of an image, diverse criteria need to be selected to assess its aesthetics. Existing works utilize pre-trained vision …

VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining

2023-03-24 · CVPR 2023 1 · Junjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu 외

Assessing the aesthetics of an image is challenging, as it is influenced by multiple factors including composition, color, style, and high-level semantics. Existing image aesthetic assessment (IAA) methods primarily rely…

DecoderLanguage ModellingVideo Quality Assessment

Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics

2026-05-20 · Yunlong Wang, Jinjin Shi, Wenbin Gao, Xuran Xu 외 arxiv

Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthetics inherently inhibit one another. Supervised Fine-Tuning (SFT) is…

Image Generation

Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms

2024-06-13 · Miaosen Zhang, Yixuan Wei, Zhen Xing, Yifei Ma 외

Modern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results in certain aspects, e.g., visual aestheti…

Retrieval

Textual Aesthetics in Large Language Models

2024-11-05 · Lingjie Jiang, Shaohan Huang, Xun Wu, Furu Wei

Image aesthetics is a crucial metric in the field of image generation. However, textual aesthetics has not been sufficiently explored. With the widespread application of large language models (LLMs), previous work has pr…

Image Generation