paper-with-me

Papers

QPT V2: Masked Image Modeling Advances Visual Scoring

2024-07-23 · Qizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu, Ming Sun, Chao Zhou, Jihong Zhu

Quality assessment and aesthetics assessment aim to evaluate the perceived quality and aesthetics of visual content. Current learning-based methods suffer greatly from the scarcity of labeled data and usually perform sub-optimally in terms of generalization. Although masked image modeling (MIM) has achieved noteworthy advancements across various high-level tasks (e.g., classification, detection etc.). In this work, we take on a novel perspective to investigate its capabilities in terms of quality- and aesthetics-awareness. To this end, we propose Quality- and aesthetics-aware pretraining (QPT V2), the first pretraining framework based on MIM that offers a unified solution to quality and aesthetics assessment. To perceive the high-level semantics and fine-grained details, pretraining data is curated. To comprehensively encompass quality- and aesthetics-related factors, degradation is introduced. To capture multi-scale quality and aesthetic information, model structure is modified. Extensive experimental results on 11 downstream benchmarks clearly show the superior performance of QPT V2 in comparison with current state-of-the-art approaches and other pretraining paradigms. Code and models will be released at \url{https://github.com/KeiChiTse/QPT-V2}.

📄 PDF Abstract BibTeX arXiv:2407.16541

Code (1)

keichitse/qpt-v2 공식 구현

Methods 이 논문이 사용한 방법론

QPT 설명 없음
MIM 설명 없음

Similar Papers 제목 키워드 기반

Masked Image Modeling Advances 3D Medical Image Analysis

2022-04-25 · Zekai Chen, Devansh Agarwal, Kshitij Aggarwal, Wiem Safta 외

Recently, masked image modeling (MIM) has gained considerable attention due to its capacity to learn from vast amounts of unlabeled data and has been demonstrated to be effective on a wide variety of vision tasks involvi…

Contrastive LearningDecoderImage SegmentationMedical Image Analysis+3

ResiComp: Loss-Resilient Image Compression via Dual-Functional Masked Visual Token Modeling

2025-02-15 · Sixian Wang, Jincheng Dai, Xiaoqi Qin, Ke Yang 외

Recent advancements in neural image codecs (NICs) are of significant compression performance, but limited attention has been paid to their error resilience. These resulting NICs tend to be sensitive to packet losses, whi…

Image CompressionMissing ValuesPacket Loss Concealment

NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining

2026-03-03 · Liang Zeng, Valerio Marsocci, Wufan Zhao, Andrea Nascetti 외 arxiv

Masked Image Modeling has been one of the most popular self-supervised learning paradigms to learn representations from large-scale, unlabeled Earth Observation images. While incorporating multi-modal and multi-temporal …

Self-Supervised Learning

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

2023-03-01 · Yuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang 외

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and …

Document Image Classificationimage-classificationImage ClassificationLanguage Modeling+4

VL-BEiT: Generative Vision-Language Pretraining

2022-06-02 · Hangbo Bao, Wenhui Wang, Li Dong, Furu Wei

We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and …

image-classificationImage ClassificationImage-text RetrievalLanguage Modeling+9