paper-with-me

홈 › Papers

Characterizing Universal Object Representations Across Vision Models

2026-05-13 · Florian P. Mahner, Johannes Roth, Ka Chun Lam, Michael F. Bonner, Francisco Pereira, Martin N. Hebart arxiv

Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. However, what remains unknown is which visual properties models actually converge on and which factors may underlie this convergence. To address this, we decompose the object similarity structure of 162 diverse vision models into a small set of non-negative dimensions. To determine universal versus model-specific dimensions, we then estimate how often each dimension reappears across models. In contrast to model-specific dimensions, universal dimensions are more interpretable and more strongly driven by conceptual image properties, indicating the relevance of interpretability and semantic content as implicit factors driving universality across models. Differences in architecture, objective function, training data, model size, and model performance do not explain the emergence of universal dimensions. However, models with more universal dimensions also better predict macaque IT activity and human similarity judgments, suggesting that universality reflects representations relevant to biological vision. These findings have important implications for understanding the emergent representations underlying deep neural network models and their alignment with biological vision.

📄 PDF Abstract BibTeX arXiv:2605.13675

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Categoroids: Universal Conditional Independence

2022-08-23 · Sridhar Mahadevan

Conditional independence has been widely used in AI, causal inference, machine learning, and statistics. We introduce categoroids, an algebraic structure for characterizing universal properties of conditional independenc…

Causal Inference

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

2020-06-30 · Fei Yu, Jiji Tang, Weichong Yin, Yu Sun 외

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic co…

AttributePredictionReferring Expression ComprehensionSentence+1

Universal dimensions of visual representation

2024-08-23 · Zirui Chen, Michael F. Bonner

Do neural network models of vision learn brain-aligned representations because they share architectural constraints and task objectives with biological vision or because they learn universal features of natural image pro…

Characterizing the temporal dynamics of universal speech representations for generalizable deepfake detection

2023-09-15 · Yi Zhu, Saurabh Powar, Tiago H. Falk

Existing deepfake speech detection systems lack generalizability to unseen attacks (i.e., samples generated by generative algorithms not seen during training). Recent studies have explored the use of universal speech rep…

DeepFake DetectionFace Swapping

Universal-Prototype Enhancing for Few-Shot Object Detection

2021-03-01 · ICCV 2021 10 · Aming Wu, Yahong Han, Linchao Zhu, Yi Yang

Few-shot object detection (FSOD) aims to strengthen the performance of novel object detection with few labeled samples. To alleviate the constraint of few samples, enhancing the generalization ability of learned features…

Few-Shot Object DetectionMeta-LearningNovel Object DetectionObject+2