paper-with-me

홈 › Papers

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

2026-04-17 · Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li arxiv

Always-on contextual AI runs language-aligned vision foundation models (VFMs) on edge devices, where the on-device model is the dominant continuous compute cost under strict latency and power limits. Due to an observed low-frequency shift in scene context and its relevant vocabulary, we present AdaDINO, an adaptive framework that makes on-device VFM inference efficient by matching execution to the current scene and task. We build on a known phenomenon, that the accuracy drop of shrinking model sizes depends on the task, and turn it into task-level adaptive execution. AdaDINO integrates neural architecture search (NAS) into a language-aligned VFM backbone distilled from DINOv2, training a single family of subnets for efficient execution during runtime. A multimodal large language model (LLM) on the cloud, invoked at low frequency, refines the candidate class set from scene context, while a learned selector activates the least-cost subnet predicted to retain a target fraction of accuracy. With the backbone and semantic pipeline held fixed, learned selection alone reduces average compute by $37\%$ over the best fixed subnet at equal segmentation accuracy. Across zero-shot classification and open-vocabulary segmentation, AdaDINO establishes a strong accuracy-efficiency frontier, improving over evaluated models of comparable sizes by up to $7.9\%$ in acc@1 on IN1K and $5.2\%$ mIoU on ADE20K, and reducing average FLOPs by up to $74.9\%$ at similar accuracy.

📄 PDF Abstract BibTeX arXiv:2604.15622

Code (0)

등록된 구현이 없습니다.

Tasks

Neural Architecture Search

Similar Papers 제목 키워드 기반

DinoComplete: 3D Shape Completion with Distilled Semantic Priors and State Space Models

2026-05-26 · Furkan Mert Algan, Eckehard Steinbach arxiv

3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insufficient for inferring missing structure. We present DinoComplete, a d…

X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning

2026-01-16 · Maanping Shao, Feihong Zhang, Gu Zhang, Baiye Cheng 외 arxiv

Visuomotor policies often leverage large pre-trained Vision Transformers (ViTs) for their powerful generalization capabilities. However, their significant data requirements present a major challenge in the data-scarce co…

Knowledge Distillation

BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning

2026-04-30 · Yizhou Wu, Shansong Wang, Yuheng Li, Mojtaba Safari 외 arxiv

Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data. Here we show that a single self-supervised represe…

Self-Supervised LearningRepresentation LearningTumor SegmentationAge Estimation

LEAP: Layer-skipping Efficiency via Adaptive Progression for Vision Transformer Distillation

2026-06-17 · Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou 외 arxiv

Vision Foundation Models (VFMs) with Vision Transformer (ViT) backbones, such as DINOv2, have become essential for downstream tasks like object recognition and semantic segmentation. The immense computational requirement…

Knowledge DistillationSemantic SegmentationObject Recognition

TinySSL: Distilled Self-Supervised Pretraining for Sub-Megabyte MCU Models

2026-05-07 · Bibin Wilson arxiv

Self-supervised learning (SSL) has transformed representation learning for large models, yet remains unexplored for microcontroller (MCU)-class models with fewer than 500K parameters. We identify three obstacles at this …

Self-Supervised LearningRepresentation Learning