paper-with-me

홈 › Papers

Efficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs

2026-07-02 · Altay Toktassyn, Jurn-Gyu Park arxiv

Modern pretrained vision models achieve strong accuracy but demand substantial GPU memory for fine-tuning, making edge deployment impractical. This paper compares five parameter-efficient fine-tuning (PEFT) methods (Full FT, LoRA, AdaLoRA, QLoRA, BitFit) on Transformers- (ViT-Small, TinyViT) and Mamba-based vision backbones (Vim-Small, MambaVision-T) under an on-device VRAM budget (e.g., 2 GB), together with three gradient-checkpointing strategies (none, static, and a proposed memory-budget-aware adaptive algorithm); and we evaluate three families of foundation-model baselines: zero-shot contrastive vision language models (OpenCLIP, SigLIP), self-supervised vision backbones with lightweight evaluation protocols (DINOv2), and autoregressive VLMs for prompt-based classification (PaliGemma, MobileVLM, SmolVLM). Experiments on CIFAR-100 and DTD report accuracy, training time, energy, and the NetScore family of multi-objective metrics, which we extend with two deployment-aware variants. QLoRA and BitFit cut energy 20-30% at a 1-2% accuracy cost; the adaptive algorithm reduces peak memory 43-79% with 9-30% energy overhead. DINOv2 surpasses fine-tuned models on CIFAR-100 (0.917 vs. 0.897) at a fraction of the energy, while small autoregressive VLMs remain uncompetitive.

📄 PDF Abstract BibTeX arXiv:2607.02158

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Results from the Paper

RankTaskDatasetModelMetrics
#3 Classification CIFAR-100 Efficient PEFT Methods with Adaptive Che Accuracy: 79

Similar Papers 제목 키워드 기반

Adaptive Parameter Selection for Tuning Vision-Language Models

2025-01-01 · CVPR 2025 1 · Yi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min Hu

Vision-language models (VLMs) like CLIP have been widely used in various specific tasks.Parameter-efficient fine-tuning (PEFT) methods, such as prompt and adapter tuning,have become key techniques for adapting these …

Few-Shot Learning

DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images

2026-07-01 · Ke Wu, Yanan Zhang, Yingjie Gao, Wenhao Li 외 arxiv

Object detection for Unmanned Aerial Vehicles (UAVs) working in open and dynamic environments is a highly challenging task. While Vision-Language Models (VLMs) have offered a powerful solution for universal object detect…

parameter-efficient fine-tuningObject Detection

Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models

2024-04-25 · Jiawei Chen, Dingkang Yang, Yue Jiang, Mingcheng Li 외

In the realm of Medical Visual Language Models (Med-VLMs), the quest for universal efficient fine-tuning mechanisms remains paramount, especially given researchers in interdisciplinary fields are often extremely short of…

Medical Visual Question Answeringparameter-efficient fine-tuningQuestion AnsweringVisual Question Answering

PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models

2025-04-19 · Nusrat Jahan Prottasha, Upama Roy Chowdhury, Shetu Mohanto, Tasfia Nuzhat 외

Large models such as Large Language Models (LLMs) and Vision Language Models (VLMs) have transformed artificial intelligence, powering applications in natural language processing, computer vision, and multimodal learning…

Domain AdaptationFederated Learningparameter-efficient fine-tuning

DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation

2025-09-25 · Ved Umrajkar arxiv

Vision-Language Models (VLMs) are foundational to critical applications like autonomous driving, medical diagnosis, and content moderation. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA enable their effi…

parameter-efficient fine-tuningAdversarial RobustnessAutonomous DrivingMedical Diagnosis