paper-with-me

홈 › Papers

Revisiting Vision Language Foundations for No-Reference Image Quality Assessment

2025-09-22 · Ankit Yadav, Ta Duc Huy, Lingqiao Liu arxiv

Large-scale vision language pre-training has recently shown promise for no-reference image-quality assessment (NR-IQA), yet the relative merits of modern Vision Transformer foundations remain poorly understood. In this work, we present the first systematic evaluation of six prominent pretrained backbones, CLIP, SigLIP2, DINOv2, DINOv3, Perception, and ResNet, for the task of No-Reference Image Quality Assessment (NR-IQA), each finetuned using an identical lightweight MLP head. Our study uncovers two previously overlooked factors: (1) SigLIP2 consistently achieves strong performance; and (2) the choice of activation function plays a surprisingly crucial role, particularly for enhancing the generalization ability of image quality assessment models. Notably, we find that simple sigmoid activations outperform commonly used ReLU and GELU on several benchmarks. Motivated by this finding, we introduce a learnable activation selection mechanism that adaptively determines the nonlinearity for each channel, eliminating the need for manual activation design, and achieving new state-of-the-art SRCC on CLIVE, KADID10K, and AGIQA3K. Extensive ablations confirm the benefits across architectures and regimes, establishing strong, resource-efficient NR-IQA baselines.

📄 PDF Abstract BibTeX arXiv:2509.17374

Code (0)

등록된 구현이 없습니다.

Tasks

No-Reference Image Quality Assessment

Similar Papers 제목 키워드 기반

Revisiting Shadow Detection from a Vision-Language Perspective

2026-05-12 · Yonghui Wang, Shaokai Liu, Wengang Zhou, Hao Feng 외 arxiv

Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to distinguish shadows from non-shadow regions. However, this formulation …

Shadow Detection

From Stability to Inconsistency: A Study of Moral Preferences in LLMs

2025-04-08 · Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat, Maria Angelica Martinez

As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introduce a Moral Foundations LLM dataset (MFD…

MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation

2024-11-07 · Han Yang, Sotiris Anagnostidis, Enis Simsar, Thomas Hofmann

We propose MegaPortrait. It's an innovative system for creating personalized portrait images in computer vision. It has three modules: Identity Net, Shading Net, and Harmonization Net. Identity Net generates learned iden…

A Reference Architecture for Designing Foundation Model based Systems

2023-04-13 · Qinghua Lu, Liming Zhu, Xiwei Xu, Zhenchang Xing 외

The release of ChatGPT, Gemini, and other large language model has drawn huge interests on foundations models. There is a broad consensus that foundations models will be the fundamental building blocks for future AI syst…

Language ModelingLanguage ModellingLarge Language Model

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning

2025-10-18 · Xiaojun Guo, Runyu Zhou, Yifei Wang, Qi Zhang 외 arxiv

Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often fail to utilize visual evidence adequately, either depending on linguistic priors…

Self-Supervised LearningReinforcement LearningGraph Learning