paper-with-me

홈 › Papers

Combining pre-trained Vision Transformers and CIDER for Out Of Domain Detection

2023-09-06 · Grégor Jouet, Clément Duhart, Francis Rousseaux, Julio Laborde, Cyril de Runz

Out-of-domain (OOD) detection is a crucial component in industrial applications as it helps identify when a model encounters inputs that are outside the training distribution. Most industrial pipelines rely on pre-trained models for downstream tasks such as CNN or Vision Transformers. This paper investigates the performance of those models on the task of out-of-domain detection. Our experiments demonstrate that pre-trained transformers models achieve higher detection performance out of the box. Furthermore, we show that pre-trained ViT and CNNs can be combined with refinement methods such as CIDER to improve their OOD detection performance even more. Our results suggest that transformers are a promising approach for OOD detection and set a stronger baseline for this task in many contexts

📄 PDF Abstract BibTeX arXiv:2309.03047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ACORT: A Compact Object Relation Transformer for Parameter Efficient Image Captioning

2022-02-11 · Jia Huei Tan, Ying Hua Tan, Chee Seng Chan, Joon Huang Chuah

Recent research that applies Transformer-based architectures to image captioning has resulted in state-of-the-art image captioning performance, capitalising on the success of Transformers on natural language tasks. Unfor…

Image CaptioningRelation

Physics-Based Benchmarking Metrics for Multimodal Synthetic Images

2025-11-19 · Kishor Datta Gupta, Marufa Kamal, Md. Mahfuzur Rahman, Fahad Rahman 외 arxiv

Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, …

Object Detection

AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

2026-03-10 · Nguyen Anh Tuong, Phan Ba Duc, Nguyen Trung Quoc, Tran Dac Thinh 외 arxiv

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work …

Visual Question AnsweringRepresentation LearningMachine TranslationImage Captioning

Harmonic Stability Analysis of Microgrids with Converter-Interfaced Distributed Energy Resources, Part I: Modelling and Theoretical Foundations

2024-08-12 · Johanna Kristin Maria Becker, Andreas Martin Kettner, Mario Paolone

This paper proposes a method for the Harmonic Stability Assessment (HSA) of power systems with a high share of Converter-Interfaced Distributed Energy Resources (CIDERs). To this end, the Harmonic State-Space (HSS) model…

DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation

2024-08-01 · Rakshith Subramanyam, Kowshik Thopalli, Vivek Narayanaswamy, Jayaraman J. Thiagarajan

Reliably detecting when a deployed machine learning model is likely to fail on a given input is crucial for ensuring safe operation. In this work, we propose DECIDER (Debiasing Classifiers to Identify Errors Reliably), a…

Attributeimage-classificationImage Classificationmodel