paper-with-me

홈 › Papers

Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms

2024-06-13 · Miaosen Zhang, Yixuan Wei, Zhen Xing, Yifei Ma, Zuxuan Wu, Ji Li, Zheng Zhang, Qi Dai, Chong Luo, Xin Geng, Baining Guo

Modern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results in certain aspects, e.g., visual aesthetic, preferred style, and responsibility. In this paper, we target the realm of visual aesthetics and aim to align vision models with human aesthetic standards in a retrieval system. Advanced retrieval systems usually adopt a cascade of aesthetic models as re-rankers or filters, which are limited to low-level features like saturation and perform poorly when stylistic, cultural or knowledge contexts are involved. We find that utilizing the reasoning ability of large language models (LLMs) to rephrase the search query and extend the aesthetic expectations can make up for this shortcoming. Based on the above findings, we propose a preference-based reinforcement learning method that fine-tunes the vision models to distill the knowledge from both LLMs reasoning and the aesthetic models to better align the vision models with human aesthetics. Meanwhile, with rare benchmarks designed for evaluating retrieval systems, we leverage large multi-modality model (LMM) to evaluate the aesthetic performance with their strong abilities. As aesthetic assessment is one of the most subjective tasks, to validate the robustness of LMM, we further propose a novel dataset named HPIR to benchmark the alignment with human aesthetics. Experiments demonstrate that our method significantly enhances the aesthetic behaviors of the vision models, under several metrics. We believe the proposed algorithm can be a general practice for aligning vision models with human values.

📄 PDF Abstract BibTeX arXiv:2406.09397

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

When Does Perceptual Alignment Benefit Vision Representations?

2024-10-14 · Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Y. Tamir 외

Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly…

Depth EstimationImage GenerationInductive BiasRetrieval+1

Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics

2026-05-20 · Yunlong Wang, Jinjin Shi, Wenbin Gao, Xuran Xu 외 arxiv

Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthetics inherently inhibit one another. Supervised Fine-Tuning (SFT) is…

Image Generation

AesRM: Improving Video Aesthetics with Expert-Level Feedback

2026-04-30 · Yujin Han, Yujie Wei, Yefei He, Xinyu Liu 외 arxiv

Despite rapid advances in photorealistic video generation, real-world applications such as filmmaking require video aesthetics, e.g., harmonious colors and cinematic lighting, beyond visual fidelity. Prior work on visual…

Video Generation

PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery

2024-09-20 · Jicheol Park, Dongwon Kim, Boseung Jeong, Suha Kwak

Text-based person search, employing free-form text queries to identify individuals within a vast image collection, presents a unique challenge in aligning visual and textual representations, particularly at the human par…

Person SearchRetrievalText based Person Search

VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining

2023-03-24 · CVPR 2023 1 · Junjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu 외

Assessing the aesthetics of an image is challenging, as it is influenced by multiple factors including composition, color, style, and high-level semantics. Existing image aesthetic assessment (IAA) methods primarily rely…

DecoderLanguage ModellingVideo Quality Assessment