paper-with-me

홈 › Papers

Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection

2024-08-28 · ICCV 2023 1 · Jinglun Li, Xinyu Zhou, Pinxue Guo, Yixuan Sun, Yiwen Huang, Weifeng Ge, Wenqiang Zhang

Detecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data from in-distribution data through joint representation learning and statistical modeling. We learn a mixture of Gaussian models for each in-distribution category. There are many Gaussian mixture models to model different visual categories. With these Gaussian models, we design an in-distribution score function by aggregating multiple Mahalanobis-based metrics. We don't use any auxiliary outlier data as training samples, which may hurt the generalization ability of out-of-distribution detection algorithms. We split the ImageNet-1k dataset into ten folds randomly. We use one fold as the in-distribution dataset and the others as out-of-distribution datasets to evaluate the proposed method. We also conduct experiments on seven popular benchmarks, including CIFAR, iNaturalist, SUN, Places, Textures, ImageNet-O, and OpenImage-O. Extensive experiments indicate that the proposed method outperforms state-of-the-art algorithms clearly. Meanwhile, we find that our visual representation has a competitive performance when compared with features learned by classical methods. These results demonstrate that the proposed method hasn't weakened the discriminative ability of visual recognition models and keeps high efficiency in detecting out-of-distribution samples.

📄 PDF Abstract BibTeX arXiv:2408.15580

Code (0)

등록된 구현이 없습니다.

Tasks

Density EstimationOut-of-Distribution DetectionRepresentation Learning

Similar Papers 제목 키워드 기반

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

2026-06-05 · Zhe Yang, Ruyi Zhang, Hongtao Chen, Wenrui Li 외 arxiv

Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint …

audio-visual event localization

S-JEA: Stacked Joint Embedding Architectures for Self-Supervised Visual Representation Learning

2023-05-19 · Alžběta Manová, Aiden Durrant, Georgios Leontidis

The recent emergence of Self-Supervised Learning (SSL) as a fundamental paradigm for learning image representations has, and continues to, demonstrate high empirical success in a variety of tasks. However, most SSL appro…

Representation LearningSelf-Supervised Learning

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

2026-02-28 · Hulingxiao He, Zhi Tan, Yuxin Peng arxiv

A high-performing, general-purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist.…

Fine-Grained Visual RecognitionContrastive Learning

Hierarchical Disentangle Network for Object Representation Learning

2019-09-25 · Shishi Qiao, Ruiping Wang, Shiguang Shan, Xilin Chen

An object can be described as the combination of primary visual attributes. Disentangling such underlying primitives is the long objective of representation learning. It is observed that categories have the natural multi…

DecoderDisentanglementGenerative Adversarial NetworkObject+1

Hierarchical Attention Fusion of Visual and Textual Representations for Cross-Domain Sequential Recommendation

2025-04-21 · Wangyu Wu, Zhenhong Chen, Siqi Song, Xianglin Qiua 외

Cross-Domain Sequential Recommendation (CDSR) predicts user behavior by leveraging historical interactions across multiple domains, focusing on modeling cross-domain preferences through intra- and inter-sequence item rel…

Decision MakingSequential Decision MakingSequential Recommendation