paper-with-me

홈 › Papers

Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions

2024-07-23 · NeurIPS 2023 11 · Kai Liu, Zhihang Fu, Chao Chen, Sheng Jin, Ze Chen, Mingyuan Tao, Rongxin Jiang, Jieping Ye

The key to OOD detection has two aspects: generalized feature representation and precise category description. Recently, vision-language models such as CLIP provide significant advances in both two issues, but constructing precise category descriptions is still in its infancy due to the absence of unseen categories. This work introduces two hierarchical contexts, namely perceptual context and spurious context, to carefully describe the precise category boundary through automatic prompt tuning. Specifically, perceptual contexts perceive the inter-category difference (e.g., cats vs apples) for current classification tasks, while spurious contexts further identify spurious (similar but exactly not) OOD samples for every single category (e.g., cats vs panthers, apples vs peaches). The two contexts hierarchically construct the precise description for a certain category, which is, first roughly classifying a sample to the predicted category and then delicately identifying whether it is truly an ID sample or actually OOD. Moreover, the precise descriptions for those categories within the vision-language framework present a novel application: CATegory-EXtensible OOD detection (CATEX). One can efficiently extend the set of recognizable categories by simply merging the hierarchical contexts learned under different sub-task settings. And extensive experiments are conducted to demonstrate CATEX's effectiveness, robustness, and category-extensibility. For instance, CATEX consistently surpasses the rivals by a large margin with several protocols on the challenging ImageNet-1K dataset. In addition, we offer new insights on how to efficiently scale up the prompt engineering in vision-language models to recognize thousands of object categories, as well as how to incorporate large language models (like GPT-3) to boost zero-shot applications. Code is publicly available at https://github.com/alibaba/catex.

📄 PDF Abstract BibTeX arXiv:2407.16725

Code (1)

alibaba/catex 공식 구현 pytorch

Tasks

Out-of-Distribution DetectionPrompt Engineering

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Chat-Capsule: A Hierarchical Capsule for Dialog-level Emotion Analysis

2022-03-23 · Yequan Wang, Xuying Meng, Yiyi Liu, Aixin Sun 외

Many studies on dialog emotion analysis focus on utterance-level emotion only. These models hence are not optimized for dialog-level emotion detection, i.e. to predict the emotion category of a dialog as a whole. More im…

Emotion Recognition

Hierarchical Adaptive Structural SVM for Domain Adaptation

2014-08-22 · Jiaolong Xu, Sebastian Ramos, David Vazquez, Antonio M. Lopez

A key topic in classification is the accuracy loss produced when the data distribution in the training (source) domain differs from that in the testing (target) domain. This is being recognized as a very relevant problem…

Domain AdaptationGeneral Classificationimage-classificationImage Classification+4

HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images

2026-04-20 · Pourya Shamsolmoali, Masoumeh Zareapoor, Michael Felsberg, Nick Pears 외 arxiv

Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resolution, scene composition, and semantic label coverage. Differences i…

Object Detection In Aerial Images

Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

2020-01-09 · CVPR 2020 6 · Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman, Yong Jae Lee 외

Existing models often leverage co-occurrences between objects and their context to improve recognition accuracy. However, strongly relying on context risks a model's generalizability, especially when typical co-occurrenc…

Attribute

Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection

2024-08-28 · ICCV 2023 1 · Jinglun Li, Xinyu Zhou, Pinxue Guo, Yixuan Sun 외

Detecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data …

Density EstimationOut-of-Distribution DetectionRepresentation Learning