Learning Model Representations Using Publicly Available Model Hubs
The weights of neural networks have emerged as a novel data modality, giving rise to the field of weight space learning. A central challenge in this area is that learning meaningful representations of weights typically requires large, carefully constructed collections of trained models, typically referred to as model zoos. These model zoos are often trained ad-hoc, requiring large computational resources, constraining the learned weight space representations in scale and flexibility. In this work, we drop this requirement by training a weight space learning backbone on arbitrary models downloaded from large, unstructured model repositories such as Hugging Face. Unlike curated model zoos, these repositories contain highly heterogeneous models: they vary in architecture and dataset, and are largely undocumented. To address the methodological challenges posed by such heterogeneity, we propose a new weight space backbone designed to handle unstructured model populations. We demonstrate that weight space representations trained on models from Hugging Face achieve strong performance, often outperforming backbones trained on laboratory-generated model zoos. Finally, we show that the diversity of the model weights in our training set allows our weight space model to generalize to unseen data modalities. By demonstrating that high-quality weight space representations can be learned in the wild, we show that curated model zoos are not indispensable, thereby overcoming a strong limitation currently faced by the weight space learning community.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval
Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that al…
Cross-Modal RetrievalRetrievalFairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers
Multimodal diffusion transformers (MM-DiTs) have emerged as the prevalent backbone for modern text-to-image generation systems. However, they exhibit critical alignment vulnerabilities, systematically manifesting severe …
Text-to-Image GenerationMode substitution induced by electric mobility hubs: results from Amsterdam
Electric mobility hubs (eHUBS) are locations where multiple shared electric modes including electric cars and e-bikes are available. To assess their potential to reduce private car use, it is important to investigate to …
Hubs and Hyperspheres: Reducing Hubness and Improving Transductive Few-shot Learning with Hyperspherical Embeddings
Distance-based classification is frequently used in transductive few-shot learning (FSL). However, due to the high-dimensionality of image representations, FSL classifiers are prone to suffer from the hubness problem, wh…
Few-Shot LearningDifferentially Categorized Structural Connectome Hubs are Involved in Differential Microstructural Basis and Functional Implications and Contribute to Individual Identification
Human brain structural networks contain sets of centrally embedded hub regions that enable efficient information communication. However, it remains largely unknown about categories of structural brain hubs and their micr…
Diffusion MRI