paper-with-me

Papers

Exploring CLIP for Assessing the Look and Feel of Images

2022-07-25 · Jianyi Wang, Kelvin C. K. Chan, Chen Change Loy

Measuring the perception of visual content is a long-standing problem in computer vision. Many mathematical models have been developed to evaluate the look or quality of an image. Despite the effectiveness of such tools in quantifying degradations such as noise and blurriness levels, such quantification is loosely coupled with human language. When it comes to more abstract perception about the feel of visual content, existing methods can only rely on supervised models that are explicitly trained with labeled data collected via laborious user study. In this paper, we go beyond the conventional paradigms by exploring the rich visual language prior encapsulated in Contrastive Language-Image Pre-training (CLIP) models for assessing both the quality perception (look) and abstract perception (feel) of images in a zero-shot manner. In particular, we discuss effective prompt designs and show an effective prompt pairing strategy to harness the prior. We also provide extensive experiments on controlled datasets and Image Quality Assessment (IQA) benchmarks. Our results show that CLIP captures meaningful priors that generalize well to different perceptual assessments. Code is avaliable at https://github.com/IceClear/CLIP-IQA.

📄 PDF Abstract BibTeX arXiv:2207.12396

Code (1)

iceclear/clip-iqa 공식 구현 pytorch

Tasks

Image Quality AssessmentNo-Reference Image Quality AssessmentVideo Quality Assessment

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation

2022-02-24 · Peter Schaldenbrand, Zhixuan Liu, Jean Oh

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control o…

Style TransferTranslation

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

2025-11-27 · Jiaying Hong, Ting Zhu, Thanet Markchom, Huizhi Liang arxiv

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels tha…

Audio GenerationMusic Generation

Connecting Look and Feel: Associating the visual and tactile properties of physical materials

2017-04-12 · CVPR 2017 7 · Wenzhen Yuan, Shaoxiong Wang, Siyuan Dong, Edward Adelson

For machines to interact with the physical world, they must understand the physical properties of objects and materials they encounter. We use fabrics as an example of a deformable material with a rich set of mechanical …

CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human Feelings

2023-11-13 · Yachun Mi, Yu Li, Yan Shu, Chen Hui 외

Video Quality Assessment (VQA) aims to simulate the process of perceiving video quality by the human visual system (HVS). The judgments made by HVS are always influenced by human subjective feelings. However, most of the…

Video Quality AssessmentVisual Question Answering (VQA)

Look into Person: Self-supervised Structure-sensitive Learning and A New Benchmark for Human Parsing

2017-03-16 · CVPR 2017 7 · Ke Gong, Xiaodan Liang, Dongyu Zhang, Xiaohui Shen 외

Human parsing has recently attracted a lot of research interests due to its huge application potentials. However existing datasets have limited number of images and annotations, and lack the variety of human appearances …

Human ParsingSelf-Supervised LearningSemantic Segmentation