paper-with-me

Papers

Aesthetic Image Captioning with Saliency Enhanced MLLMs

2025-09-04 · Yilin Tao, Jiashui Huang, Huaze Xu, Ling Shao arxiv

Aesthetic Image Captioning (AIC) aims to generate textual descriptions of image aesthetics, becoming a key research direction in the field of computational aesthetics. In recent years, pretrained Multimodal Large Language Models (MLLMs) have advanced rapidly, leading to a significant increase in image aesthetics research that integrates both visual and textual modalities. However, most existing studies on image aesthetics primarily focus on predicting aesthetic ratings and have shown limited application in AIC. Existing AIC works leveraging MLLMs predominantly rely on fine-tuning methods without specifically adapting MLLMs to focus on target aesthetic content. To address this limitation, we propose the Aesthetic Saliency Enhanced Multimodal Large Language Model (ASE-MLLM), an end-to-end framework that explicitly incorporates aesthetic saliency into MLLMs. Within this framework, we introduce the Image Aesthetic Saliency Module (IASM), which efficiently and effectively extracts aesthetic saliency features from images. Additionally, we design IAS-ViT as the image encoder for MLLMs, this module fuses aesthetic saliency features with original image features via a cross-attention mechanism. To the best of our knowledge, ASE-MLLM is the first framework to integrate image aesthetic saliency into MLLMs specifically for AIC tasks. Extensive experiments demonstrated that our approach significantly outperformed traditional methods and generic MLLMs on current mainstream AIC benchmarks, achieving state-of-the-art (SOTA) performance.

📄 PDF Abstract BibTeX arXiv:2509.04378

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

2025-12-25 · Shuo Cao, Jiayang Li, Xiaohui Li, Yuandong Pu 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image f…

Visual Question AnsweringText-to-Image GenerationVisual Grounding

AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception

2024-01-16 · Yipo Huang, Quan Yuan, Xiangfei Sheng, Zhichao Yang 외

With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in re…

MLLM Evaluation: Aesthetics

Aesthetic Critiques Generation for Photos

2017-10-01 · ICCV 2017 10 · Kuang-Yu Chang, Kung-Hung Lu, Chu-Song Chen

It is said that a picture is worth a thousand words. Thus, there are various ways to describe an image, especially in aesthetic quality analysis. Although aesthetic quality assessment has generated a great deal of intere…

Image Captioning

Aesthetically Relevant Image Captioning

2022-11-25 · Zhipeng Zhong, Fei Zhou, Guoping Qiu

Image aesthetic quality assessment (AQA) aims to assign numerical aesthetic ratings to images whilst image aesthetic captioning (IAC) aims to generate textual descriptions of the aesthetic aspects of images. In this pape…

Image CaptioningSentence

AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception

2024-04-15 · Yipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan 외

The highly abstract nature of image aesthetics perception (IAP) poses significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerba…