Feature Representation Transferring to Lightweight Models via Perception Coherence
In this paper, we propose a method for transferring feature representation to lightweight student models from larger teacher models. We mathematically define a new notion called \textit{perception coherence}. Based on this notion, we propose a loss function, which takes into account the dissimilarities between data points in feature space through their ranking. At a high level, by minimizing this loss function, the student model learns to mimic how the teacher model \textit{perceives} inputs. More precisely, our method is motivated by the fact that the representational capacity of the student model is weaker than the teacher model. Hence, we aim to develop a new method allowing for a better relaxation. This means that, the student model does not need to preserve the absolute geometry of the teacher one, while preserving global coherence through dissimilarity ranking. Our theoretical insights provide a probabilistic perspective on the process of feature representation transfer. Our experiments results show that our method outperforms or achieves on-par performance compared to strong baseline methods for representation transferring.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
All-in-One Transferring Image Compression from Human Perception to Multi-Machine Perception
Efficiently transferring Learned Image Compression (LIC) model from human perception to machine perception is an emerging challenge in vision-centric representation learning. Existing approaches typically adapt LIC to do…
AllImage CompressionRepresentation LearningTransferring Modality-Aware Pedestrian Attentive Learning for Visible-Infrared Person Re-identification
Visible-infrared person re-identification (VI-ReID) aims to search the same pedestrian of interest across visible and infrared modalities. Existing models mainly focus on compensating for modality-specific information to…
Data AugmentationPerson Re-IdentificationTransferable Representation Learning in Vision-and-Language Navigation
Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The…
Representation LearningVision and Language NavigationCSDNet: Detect Salient Object in Depth-Thermal via A Lightweight Cross Shallow and Deep Perception Network
While we enjoy the richness and informativeness of multimodal data, it also introduces interference and redundancy of information. To achieve optimal domain interpretation with limited resources, we propose CSDNet, a lig…
DescriptiveInformativenessobject-detectionObject Detection+1Temporal Coherence or Temporal Motion: Which is More Critical for Video-based Person Re-identification?
Video-based person re-identification aims to match pedestrians with the consecutive video sequences. While a rich line of work focuses solely on extracting the motion features from pedestrian videos, we show in this pape…
Person Re-IdentificationVideo-Based Person Re-Identification