paper-with-me

Papers

Beyond Compression: Quantifying Spectral Accessibility in Vision Representations

2026-06-02 · Akayou A. Kitessa, Yijun Zhao arxiv

Vision-language models map visual features into a shared embedding space through learned projection layers, yet it remains unclear how these transformations alter the structure of visual information. This study examines changes in representation through spatial-frequency accessibility, measured by the linear recoverability of band-limited Fourier energy from model representations. To isolate effects beyond dimensionality reduction, we introduce Residual Spectral Loss (RSL), which evaluates changes relative to a dimension-matched random projection baseline. To reduce confounding effects from optimization, the analysis uses pretrained models with all parameters frozen. The experimental results show consistent frequency-dependent changes in accessibility across CLIP and DINOv2 on ImageNet and MS-COCO datasets. Spectral accessibility follows a non-monotonic trajectory across depth, peaking at intermediate layers before decreasing toward the output representation. The final transformation differs across architectures: CLIP's learned projection is spectrally neutral, with changes explained by compression, whereas DINOv2's [CLS] pooling induces a structured loss across the spectrum. These findings identify intermediate layers and pooling mechanisms as primary drivers of spectral transformation in modern vision encoders.

📄 PDF Abstract BibTeX arXiv:2606.03795

Code (0)

등록된 구현이 없습니다.

Tasks

Dimensionality Reduction

Similar Papers 제목 키워드 기반

Hyperspectral Variational Autoencoders for Joint Data Compression and Component Extraction

2025-11-23 · Core Francisco Park, Manuel Perez-Carrasco, Caroline Nowlan, Cecilia Garraffo arxiv

Geostationary hyperspectral satellites generate terabytes of data daily, creating critical challenges for storage, transmission, and distribution to the scientific community. We present a variational autoencoder (VAE) ap…

Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics

2026-03-02 · Samhruth Ananthanarayanan, Ayan Sengupta, Tanmoy Chakraborty arxiv

As context windows in LLMs scale to 100K+ tokens, the key-value (KV) cache becomes the dominant memory bottleneck, with recent methods claiming 80-90% savings and minimal benchmark degradation. We argue these evaluations…

Beyond Appearances: Material Segmentation with Embedded Spectral Information from RGB-D imagery

2024-05-17 · Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops 2024 5 · Fabian Perez, Hoover Rueda-Chacón

In the realm of computer vision material segmentation of natural scenes represents a challenge driven by the complex and diverse appearances of materials. Traditional approaches often rely on RGB images which can be dece…

Material ClassificationMaterial RecognitionMaterial SegmentationScene Understanding+1

Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects

2024-11-19 · Ke Liang Xiao, Noah Marshall, Atish Agarwala, Elliot Paquette

In recent years, signSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consensus that signSGD acts to precondition op…

Face-GPS: A Comprehensive Technique for Quantifying Facial Muscle Dynamics in Videos

2024-01-11 · Juni Kim, Zhikang Dong, Pawel Polak

We introduce a novel method that combines differential geometry, kernels smoothing, and spectral analysis to quantify facial muscle activity from widely accessible video recordings, such as those captured on personal sma…