The Devil is in the Middle: Exploiting Mid-level Representations for Cross-Domain Instance Matching
Many vision problems require matching images of object instances across different domains. These include fine-grained sketch-based image retrieval (FG-SBIR) and Person Re-identification (person ReID). Existing approaches attempt to learn a joint embedding space where images from different domains can be directly compared. In most cases, this space is defined by the output of the final layer of a deep neural network (DNN), which primarily contains features of a high semantic level. In this paper, we argue that both high and mid-level features are relevant for cross-domain instance matching (CDIM). Importantly, mid-level features already exist in earlier layers of the DNN. They just need to be extracted, represented, and fused properly with the final layer. Based on this simple but powerful idea, we propose a unified framework for CDIM. Instantiating our framework for FG-SBIR and ReID, we show that our simple models can easily beat the state-of-the-art models, which are often equipped with much more elaborate architectures.
Code (0)
등록된 구현이 없습니다.
Tasks
Image RetrievalPerson Re-IdentificationRetrievalSketch-Based Image RetrievalSimilar Papers 제목 키워드 기반
Middle-level Fusion for Lightweight RGB-D Salient Object Detection
Most existing lightweight RGB-D salient object detection (SOD) models are based on two-stream structure or single-stream structure. The former one first uses two sub-networks to extract unimodal features from RGB and dep…
object-detectionObject DetectionRGB-D Salient Object DetectionSalient Object DetectionCAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
Despite the transformative impact of deep learning across multiple domains, the inherent opacity of these models has driven the development of Explainable Artificial Intelligence (XAI). Among these efforts, Concept Bottl…
Backdoor AttackExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspec…
HallucinationExploiting auto-encoders and segmentation methods for middle-level explanations of image classification systems
A central issue addressed by the rapidly growing research area of eXplainable Artificial Intelligence (XAI) is to provide methods to give explanations for the behaviours of Machine Learning (ML) non-interpretable models …
Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)image-classificationImage ClassificationA metapopulation model of the spread of the Devil Facial Tumour Disease predicts the long term collapse of its host but not its extinction
The Devil Facial Tumour Disease (DFTD), a unique case of a transmissible cancer, had a devastating effect on its host, the Tasmanian Devil. Current estimates of its density are at roughly 20% of the pre-disease state, an…