paper-with-me

Papers

Stratified Knowledge-Density Super-Network for Scalable Vision Transformers

2025-11-12 · Longhua Li, Lei Qi, Xin Geng arxiv

Training and deploying multiple vision transformer (ViT) models for different resource constraints is costly and inefficient. To address this, we propose transforming a pre-trained ViT into a stratified knowledge-density super-network, where knowledge is hierarchically organized across weights. This enables flexible extraction of sub-networks that retain maximal knowledge for varying model sizes. We introduce \textbf{W}eighted \textbf{P}CA for \textbf{A}ttention \textbf{C}ontraction (WPAC), which concentrates knowledge into a compact set of critical weights. WPAC applies token-wise weighted principal component analysis to intermediate features and injects the resulting transformation and inverse matrices into adjacent layers, preserving the original network function while enhancing knowledge compactness. To further promote stratified knowledge organization, we propose \textbf{P}rogressive \textbf{I}mportance-\textbf{A}ware \textbf{D}ropout (PIAD). PIAD progressively evaluates the importance of weight groups, updates an importance-aware dropout list, and trains the super-network under this dropout regime to promote knowledge stratification. Experiments demonstrate that WPAC outperforms existing pruning criteria in knowledge concentration, and the combination with PIAD offers a strong alternative to state-of-the-art model compression and model expansion methods.

📄 PDF Abstract BibTeX arXiv:2511.11683

Code (0)

등록된 구현이 없습니다.

Tasks

Model Compression

Similar Papers 제목 키워드 기반

Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling

2026-03-17 · Hongjian Zou, Yue Ge, Qi Ding, Yixuan Liao 외 arxiv

Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less predictable than that of text-only LLMs. Increasing model size and task …

Visual Question Answering

Cross-position Activity Recognition with Stratified Transfer Learning

2018-06-26 · Yiqiang Chen, Jindong Wang, Meiyu Huang, Han Yu

Human activity recognition aims to recognize the activities of daily living by utilizing the sensors on different body parts. However, when the labeled data from a certain body position (i.e. target domain) is missing, h…

Activity RecognitionHuman Activity RecognitionPositionTransfer Learning

Is sub-metre resolution necessary for cocoa mapping? A landscape-stratified evaluation of very high resolution imagery, decametric Earth Observation inputs, and operational products in Cote d'Ivoire

2026-07-09 · Kasimir Orlowski, Filip Sabo, Michele Meroni, Astrid Verhegghen 외 arxiv

Accurate cocoa mapping is increasingly important for deforestation monitoring, supply-chain transparency, and regulatory applications. Spatial aggregation in conventional medium-resolution Earth observation (EO) imagery …

Spatially Stratified Distillation for Heterogeneous Radar Place Recognition

2026-06-17 · Sagun Singh Shrestha, Samuel Harding, Abdelwahed Khamis, Saimunur Rahman 외 arxiv

Scalable, all-weather place recognition increasingly relies on heterogeneous radar place recognition to bridge diverse hardware platforms. A notable application is matching queries from cost-effective 4D automotive radar…

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

2026-07-05 · Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang 외 arxiv

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At t…

Reinforcement LearningSound Event Detection