paper-with-me

홈 › Papers

What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?

2026-04-13 · Koki Ryu, Hitomi Yanaka arxiv

Personalized image aesthetics assessment (PIAA) is an important research problem with practical real-world applications. While methods based on vision-language models (VLMs) are promising candidates for PIAA, it remains unclear whether they internally encode rich, multi-level aesthetic attributes required for effective personalization. In this paper, we first analyze the internal representations of VLMs to examine the presence and distribution of such aesthetic attributes, and then leverage them for lightweight, individual-level personalization without model fine-tuning. Our analysis reveals that VLMs encode diverse aesthetic attributes that propagate into the language decoder layers. Building on these representations, we demonstrate that simple linear models can perform PIAA effectively. We further analyze how aesthetic information is transferred across layers in different VLM architectures and across image domains. Our findings provide insights into how VLMs can be utilized for modeling subjective, individual aesthetic preferences. Our code is available at https://github.com/ynklab/vlm-latent-piaa.

📄 PDF Abstract BibTeX arXiv:2604.11374

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Yo'LLaVA: Your Personalized Language and Vision Assistant

2024-06-13 · Thao Nguyen, Haotian Liu, Yuheng Li, Mu Cai 외

Large Multimodal Models (LMMs) have shown remarkable capabilities across a variety of tasks (e.g., image captioning, visual question answering). While broad, their knowledge remains generic (e.g., recognizing a dog), and…

Image CaptioningQuestion AnsweringVisual Question Answering

Personalized Image Descriptions from Attention Sequences

2025-12-07 · Ruoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal 외 arxiv

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descript…

Improving Personalized Search with Regularized Low-Rank Parameter Updates

2025-06-11 · CVPR 2025 1 · Fiona Ryan, Josef Sivic, Fabian Caba Heilbron, Judy Hoffman 외

Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but al…

General KnowledgeImage RetrievalNatural Language QueriesRetrieval

"This is my unicorn, Fluffy": Personalizing frozen vision-language representations

2022-04-04 · Niv Cohen, Rinon Gal, Eli A. Meirom, Gal Chechik 외

Large Vision & Language models pretrained on web-scale data provide representations that are invaluable for numerous V&L problems. However, it is unclear how they can be used for reasoning about user-specific visual conc…

Image RetrievalRetrievalSemantic SegmentationSentence+3

Personalized Representation from Personalized Generation

2024-12-20 · Shobhita Sundaram, Julia Chae, Yonglong Tian, Sara Beery 외

Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-scarce. Recent works have successfully appl…

Contrastive LearningImage GenerationRepresentation Learning