paper-with-me

홈 › Papers

Dimensional Coactivation for Representational Consistency in Frozen Vision Foundation Models

2026-05-07 · Izaldein Al-Zyoud Abdulmotaleb El Saddik arxiv

Frozen vision foundation models do not merely extract features; they organize images through a learned coordinate system. We ask whether that coordinate system remains internally coherent within a single input. This leads to Representational Consistency: the study of whether a frozen foundation model represents one sample coherently across its semantic subregions. We introduce Dimensional Coactivation (DCA), a per-dimension instrument for measuring this coherence. DCA compares semantic regions by asking whether the same feature dimensions coactivate across them. Unlike classical similarity measures, it deliberately avoids centering, L2 normalization, and full Gram coupling. These operations are useful when comparing different models or distributions, but they are mismatched to the intra-sample setting, where the coordinate system is fixed and raw magnitude carries signal. Deepfake detection provides a natural validation task. Synthetic faces may reproduce plausible eyes, noses, and mouths while breaking the representational structure that links those regions in real faces. Using frozen DINOv3 features, DCA exposes this break: an eyes-mouth-nose fingerprint achieves 0.9106 AUC on CelebDF-v2 and 0.9289 on DFD under FF++ c23 cross-dataset transfer. The design is also sharply validated by ablation: reintroducing centering collapses CelebDF-v2 AUC to 0.459, L2 normalization reduces it to 0.862, and cross-dimension coupling reduces it to 0.478. Finally, replacing DINOv3 with FaRL collapses CelebDF-v2 AUC to 0.582. DCA therefore depends on a stable per-dimension coordinate system, not on region extraction alone. These results position DCA as an instrument for measuring intra-sample representational coherence in frozen foundation models, with deepfake detection as the first validation task.

📄 PDF Abstract BibTeX arXiv:2605.08249

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake Detection

Similar Papers 제목 키워드 기반

Muscle coactivation primes the nervous system for fast and task-dependent feedback control

2024-10-21 · Philipp Maurus, Daniel P. Armstrong, Stephen H. Scott, Tyler Cluff

Humans and other animals coactivate agonist and antagonist muscles in many motor actions. Increases in muscle coactivation are thought to leverage viscoelastic properties of skeletal muscles to provide resistance against…

Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning

2026-02-27 · Abhishek Dalvi, Vasant Honavar arxiv

Large unimodal foundation models for vision and language encode rich semantic structures, yet aligning them typically requires computationally intensive multimodal fine-tuning. Such approaches depend on large-scale param…

Image Captioning

Do Foundation Models Know Geometry? Probing Frozen Features for Continuous Physical Measurement

2026-03-06 · Yakov Pyotr Shkolnikov arxiv

Vision-language models encode continuous geometry that their text pathway fails to express: a 6,000-parameter linear probe extracts hand joint angles at 6.1 degrees MAE from frozen features, while the best text output ac…

Text Generation

MindAdapter: Few-Shot Parameter-Efficient Residual Calibration of Cross-Subject Brain-to-Visual Decoding Models

2026-05-23 · Jiaxiang Liu, Jiawei Du, Xupeng Chen, Guoqi Li 외 arxiv

Cross-subject brain-to-visual decoding remains a core challenge in brain-computer interfaces due to severe inter-individual variability that induces systematic subject-specific functional misalignment. To address this is…

Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory

2026-06-17 · Kaustubh Kapil, Kishor P. Upla arxiv

Despite the widespread adoption of Vision Transformers (ViTs) and their success across numerous computer vision applications, the fundamental understanding of their dimensional and representational geometry remains relat…