paper-with-me

홈 › Papers

ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation

2025-03-08 · Qizhen Lan, Qing Tian

Dense visual prediction tasks, such as detection and segmentation, are crucial for time-critical applications (e.g., autonomous driving and video surveillance). While deep models achieve strong performance, their efficiency remains a challenge. Knowledge distillation (KD) is an effective model compression technique, but existing feature-based KD methods rely on static, teacher-driven feature selection, failing to adapt to the student's evolving learning state or leverage dynamic student-teacher interactions. To address these limitations, we propose Adaptive student-teacher Cooperative Attention Masking for Knowledge Distillation (ACAM-KD), which introduces two key components: (1) Student-Teacher Cross-Attention Feature Fusion (STCA-FF), which adaptively integrates features from both models for a more interactive distillation process, and (2) Adaptive Spatial-Channel Masking (ASCM), which dynamically generates importance masks to enhance both spatial and channel-wise feature selection. Unlike conventional KD methods, ACAM-KD adapts to the student's evolving needs throughout the entire distillation process. Extensive experiments on multiple benchmarks validate its effectiveness. For instance, on COCO2017, ACAM-KD improves object detection performance by up to 1.4 mAP over the state-of-the-art when distilling a ResNet-50 student from a ResNet-101 teacher. For semantic segmentation on Cityscapes, it boosts mIoU by 3.09 over the baseline with DeepLabV3-MobileNetV2 as the student model.

📄 PDF Abstract BibTeX arXiv:2503.06307

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Drivingfeature selectionKnowledge DistillationModel Compressionobject-detectionObject DetectionSemantic Segmentation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

A Knowledge-Enhanced Recommendation Model with Attribute-Level Co-Attention

2020-06-18 · Deqing Yang, Zengcun Song, Lvxin Xue, Yanghua Xiao

Deep neural networks (DNNs) have been widely employed in recommender systems including incorporating attention mechanism for performance improvement. However, most of existing attention-based models only apply item-level…

AttributeKnowledge GraphsRecommendation Systems

From Cheap to Pro: A Learning-based Adaptive Camera Parameter Network for Professional-Style Imaging

2025-10-23 · Fuchen Li, Yansong Du, Wenbo Cheng, Xiaoxia Zhou 외 arxiv

Consumer-grade camera systems often struggle to maintain stable image quality under complex illumination conditions such as low light, high dynamic range, and backlighting, as well as spatial color temperature variation.…

Image Enhancement

MetaCAM: Ensemble-Based Class Activation Map

2023-07-31 · Emily Kaczmarek, Olivier X. Miguel, Alexa C. Bowie, Robin Ducharme 외

The need for clear, trustworthy explanations of deep learning model predictions is essential for high-criticality fields, such as medicine and biometric identification. Class Activation Maps (CAMs) are an increasingly po…

Doracamom: Joint 3D Detection and Occupancy Prediction with Multi-view 4D Radars and Cameras for Omnidirectional Perception

2025-01-26 · Lianqing Zheng, Jianan Liu, Runwei Guan, Long Yang 외

3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse condi…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection+1

Machine Learning-Based Classification of Jhana Advanced Concentrative Absorption Meditation (ACAM-J) using 7T fMRI

2026-02-13 · Puneet Kumar, Winson F. Z. Yang, Alakhsimar Singh, Xiaobai Li 외 arxiv

Jhana advanced concentration absorption meditation (ACAM-J) is related to profound changes in consciousness and cognitive processing, making the study of their neural correlates vital for insights into consciousness and …