paper-with-me

홈 › Papers

ERank in Latent Space as an Image-Complexity and Richness Measure

2026-07-21 · Maksim Smirnov, Grigory Kononov, Anastasiia Linich, Egor Surkov, Egor Shvetsov arxiv

We propose the effective rank (ERank) of the channel covariance of an image's deep feature map as a per-sample, label-free measure of visual richness, computed from a single forward pass through a frozen pretrained encoder. ERank counts how many decorrelated channel directions an image activates, and we characterize its properties, including its behavior under noise. Empirically, ERank orders images from plain to visually rich, correlates with codec bitrate, sharpness, and edge density, and correlates with human complexity annotations on IC9600 with $r = 0.72$. As a data-selection criterion, removing low-ERank samples improves super-resolution and removing high-ERank samples improves OCR, in both pretraining and finetuning, while selection does not help classification, segmentation, or denoising. ERank is thus a cheap richness signal, useful exactly when task difficulty is governed by input richness.

📄 PDF Abstract BibTeX arXiv:2607.19315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Traversing Latent Space using Decision Ferns

2018-12-06 · Yan Zuo, Gil Avraham, Tom Drummond

The practice of transforming raw data to a feature space so that inference can be performed in that space has been popular for many years. Recently, rapid progress in deep neural networks has given both researchers and p…

Adversarial attacks to image classification systems using evolutionary algorithms

2025-07-17 · Sergio Nesmachnow, Jamal Toutouh

Image classification currently faces significant security challenges due to adversarial attacks, which consist of intentional alterations designed to deceive classification models based on artificial intelligence. This a…

ClassificationDiversityEvolutionary AlgorithmsGenerative Adversarial Network+2

Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference

2025-07-02 · Xu Zhang, Ming Lu, Yan Chen, Zhan Ma

In recent years, compressed domain semantic inference has primarily relied on learned image coding models optimized for mean squared error (MSE). However, MSE-oriented optimization tends to yield latent spaces with limit…

Image ClassificationImage CompressionSemantic Segmentation

Semantic-Enriched Latent Visual Reasoning

2026-05-19 · Tianrun Xu, Yue Sun, Qixun Wang, Jingyi Lu 외 arxiv

Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches largely rely on visual supervision and produ…

Question AnsweringVisual Reasoning

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

2026-08-25 · Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim 외 arxiv

Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with variational autoencoders (VAEs) to enhance computational efficiency withou…

Computational Efficiency