paper-with-me

홈 › Papers

Deep Psychovisual Image Representations

2026-05-28 · Wendi Ma, Aryaman Sharma, Wei Dai, Shekhar S. Chandra arxiv

Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In contrast, deep learning-based vision models routinely extract and aggregate features using homogeneous stacks of spatial layers, rendering their decision-making processes opaque. In this paper, we propose Deep Visual Coding, a learned frequency-domain representation inspired by 1990s image codes that quantised perceptually salient frequencies, which together with complex-valued image representations produces psychovisual-style abstractions. This approach enables the first psychovisual-based deep learning framework, utilizing data-driven spectral filters that learn to encode task-relevant semantic structures within distinct frequency sub-bands. Salience analyses reveal that our psychovisual models extract highly interpretable object parts compared to the amorphous regions produced by regular Convolutional Neural Networks (CNNs). Furthermore, we find that our models are less depth dependent than CNNs for model scaling, since our complex-valued representations and learned abstractions subsume the role of the deep spatial layers. Together, these findings demonstrate that psychovisual coding provides a promising path toward more efficient and transparent vision models.

📄 PDF Abstract BibTeX arXiv:2605.29260

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Color graph based wavelet transform with perceptual information

2015-10-19 · Mohamed Malek, David Helbert, Philippe Carre

In this paper, we propose a numerical strategy to define a multiscale analysis for color and multicomponent images based on the representation of data on a graph. Our approach consists in computing the graph of an image …

DenoisingImage Restoration

Users prefer Guetzli JPEG over same-sized libjpeg

2017-03-13 · Jyrki Alakuijala, Robert Obryk, Zoltan Szabadka, Jan Wassenberg

We report on pairwise comparisons by human raters of JPEG images from libjpeg and our new Guetzli encoder. Although both files are size-matched, 75% of ratings are in favor of Guetzli. This implies the Butteraugli psycho…

A Mixed Bag of Emotions: Model, Predict, and Transfer Emotion Distributions

2015-06-01 · CVPR 2015 6 · Kuan-Chuan Peng, Tsuhan Chen, Amir Sadovnik, Andrew C. Gallagher

This paper explores two new aspects of photos and human emotions. First, we show through psychovisual studies that different people have different emotional reactions to the same image, which is a strong and novel depart…

Adaptive Blind Watermarking Using Psychovisual Image Features

2022-12-25 · Arezoo PariZanganeh, Ghazaleh Ghorbanzadeh, Zahra Nabizadeh ShahreBabak, Nader Karimi 외

With the growth of editing and sharing images through the internet, the importance of protecting the images' authorship has increased. Robust watermarking is a known approach to maintaining copyright protection. Robustne…

Guetzli: Perceptually Guided JPEG Encoder

2017-03-13 · Jyrki Alakuijala, Robert Obryk, Ostap Stoliarchuk, Zoltan Szabadka 외

Guetzli is a new JPEG encoder that aims to produce visually indistinguishable images at a lower bit-rate than other common JPEG encoders. It optimizes both the JPEG global quantization tables and the DCT coefficient valu…

Perceptual DistanceQuantization