paper-with-me

홈 › Papers

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot

2025-01-15 · Ruixiang Jiang, Changwen Chen

The rapid progress of generative art has democratized the creation of visually pleasing imagery. However, achieving genuine artistic impact - the kind that resonates with viewers on a deeper, more meaningful level - requires a sophisticated aesthetic sensibility. This sensibility involves a multi-faceted reasoning process extending beyond mere visual appeal, which is often overlooked by current computational models. This paper pioneers an approach to capture this complex process by investigating how the reasoning capabilities of Multimodal LLMs (MLLMs) can be effectively elicited for aesthetic judgment. Our analysis reveals a critical challenge: MLLMs exhibit a tendency towards hallucinations during aesthetic reasoning, characterized by subjective opinions and unsubstantiated artistic interpretations. We further demonstrate that these limitations can be overcome by employing an evidence-based, objective reasoning process, as substantiated by our proposed baseline, ArtCoT. MLLMs prompted by this principle produce multi-faceted and in-depth aesthetic reasoning that aligns significantly better with human judgment. These findings have direct applications in areas such as AI art tutoring and as reward models for generative art. Ultimately, our work paves the way for AI systems that can truly understand, appreciate, and generate artworks that align with the sensible human aesthetic standard.

📄 PDF Abstract BibTeX arXiv:2501.09012

Code (1)

songrise/mllm4art 공식 구현

Tasks

BenchmarkingHallucinationImage GenerationStyle Transfer

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM

2026-04-19 · Chun Wang, Chenfeng Wei, Chenyang Liu, Weihong Deng arxiv

Personalized image aesthetics assessment (PIAA) aims to predict an individual user's subjective rating of an image, which requires modeling user-specific aesthetic preferences. Existing methods rely on historical user ra…

Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior

2026-04-03 · Nolan Platt, Sehrish Nizamani, Alp Tural, Elif Tural 외 arxiv

Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns. We present a privacy-preserving pipeline that analyzes classroom videos to extract i…

Spatial Reasoning

Aesthetic Image Captioning with Saliency Enhanced MLLMs

2025-09-04 · Yilin Tao, Jiashui Huang, Huaze Xu, Ling Shao arxiv

Aesthetic Image Captioning (AIC) aims to generate textual descriptions of image aesthetics, becoming a key research direction in the field of computational aesthetics. In recent years, pretrained Multimodal Large Languag…

Image Captioning

Exploring Failure Cases in Multimodal Reasoning About Physical Dynamics

2024-02-24 · Sadaf Ghaffari, Nikhil Krishnaswamy

In this paper, we present an exploration of LLMs' abilities to problem solve with physical reasoning in situated environments. We construct a simple simulated environment and demonstrate examples of where, in a zero-shot…

Language ModelingLanguage ModellingMultimodal ReasoningObject+1

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning

2026-06-29 · William Orwig, Roger E. Beaty arxiv

Evaluating the originality of visual images poses enduring challenges for creativity assessment. Automated scoring using AI models has proven effective in the verbal domain, yet key questions remain about evaluating visu…