paper-with-me

홈 › Papers

Module-wise Adaptive Distillation for Multimodality Foundation Models

2023-10-06 · NeurIPS 2023 11

Pre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reducing their sizes is layerwise distillation, wherein small student models are trained to match the hidden representations of large teacher models at each layer. Motivated by our observation that certain architecture components, referred to as modules, contribute more significantly to the student's performance than others, we propose to track the contributions of individual modules by recording the loss decrement after distillation each module and choose the module with a greater contribution to distill more frequently. Such an approach can be naturally formulated as a multi-armed bandit (MAB) problem, where modules and loss decrements are considered as arms and rewards, respectively. We then develop a modified-Thompson sampling algorithm named OPTIMA to address the nonstationarity of module contributions resulting from model updating. Specifically, we leverage the observed contributions in recent history to estimate the changing contribution of each module and select modules based on these estimations to maximize the cumulative contribution. We evaluate the effectiveness of OPTIMA through distillation experiments on various multimodal understanding and image captioning tasks, using the CoCa-Large model (Yu et al., 2022) as the teacher model.

📄 PDF Abstract BibTeX arXiv:2310.04550

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningThompson Sampling

Similar Papers 제목 키워드 기반

Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition

2026-08-04 · Ping Li, Chenhao Ping, Jie Song, Mingli Song arxiv

Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models. However, existing KD methods suffer from two key limitations: 1) reliance on fixed input samples lead…

Knowledge DistillationAction Recognition

Multimodality Multi-Lead ECG Arrhythmia Classification using Self-Supervised Learning

2022-09-30 · Thinh Phan, Duc Le, Patel Brijesh, Donald Adjeroh 외

Electrocardiogram (ECG) signal is one of the most effective sources of information mainly employed for the diagnosis and prediction of cardiovascular diseases (CVDs) connected with the abnormalities in heart rhythm. Clea…

ECG ClassificationKnowledge DistillationRhythmSelf-Knowledge Distillation+3

FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation

2026-07-30 · Lifeng Zhuo, Wendi Chen, Han Xue, Shirun Tang 외 arxiv

In contact-rich manipulation, action multimodality and reactivity dominate different stages of a single episode. Before contact, multiple trajectories might be equally valid, making it important to preserve diverse actio…

MAFM^3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI

2025-11-14 · Mohammad Areeb Qazi, Munachiso S Nwadike, Ibrahim Almakky, Mohammad Yaqub 외 arxiv

Foundational models are trained on extensive datasets to capture the general trends of a domain. However, in medical imaging, the scarcity of data makes pre-training for every domain, modality, or task challenging. Inste…

Context-Aware Knowledge Distillation with Adaptive Weighting for Image Classification

2025-08-30 · Zhengda Li arxiv

Knowledge distillation (KD) is a widely used technique to transfer knowledge from a large teacher network to a smaller student model. Traditional KD uses a fixed balancing factor alpha as a hyperparameter to combine the …

Knowledge DistillationImage Classification