Forte : Finding Outliers with Representation Typicality Estimation
Generative models can now produce photorealistic synthetic data which is virtually indistinguishable from the real data used to train it. This is a significant evolution over previous models which could produce reasonable facsimiles of the training data, but ones which could be visually distinguished from the training data by human evaluation. Recent work on OOD detection has raised doubts that generative model likelihoods are optimal OOD detectors due to issues involving likelihood misestimation, entropy in the generative process, and typicality. We speculate that generative OOD detectors also failed because their models focused on the pixels rather than the semantic content of the data, leading to failures in near-OOD cases where the pixels may be similar but the information content is significantly different. We hypothesize that estimating typical sets using self-supervised learners leads to better OOD detectors. We introduce a novel approach that leverages representation learning, and informative summary statistics based on manifold estimation, to address all of the aforementioned issues. Our method outperforms other unsupervised approaches and achieves state-of-the art performance on well-established challenging benchmarks, and new synthetic data detection tasks.
Code (1)
Tasks
Out-of-Distribution DetectionRepresentation LearningSimilar Papers 제목 키워드 기반
How Well Do Deep Learning Models Capture Human Concepts? The Case of the Typicality Effect
How well do representations learned by ML models align with those of humans? Here, we consider concept representations learned by deep learning models and evaluate whether they show a fundamental behavioral signature of …
Language ModelingLanguage ModellingBenchmarking VLMs' Reasoning About Persuasive Atypical Images
Vision language models (VLMs) have shown strong zero-shot generalization across various tasks, especially when integrated with large language models (LLMs). However, their ability to comprehend rhetorical and persuasive …
BenchmarkingObject RecognitionZero-shot GeneralizationStudy of Robust Direction Finding Based on Joint Sparse Representation
Standard Direction of Arrival (DOA) estimation methods are typically derived based on the Gaussian noise assumption, making them highly sensitive to outliers. Therefore, in the presence of impulsive noise, the performanc…
Modeling the human lexicon under temperature variations: linguistic factors, diversity and typicality in LLM word associations
Large language models (LLMs) achieve impressive results in terms of fluency in text generation, yet the nature of their linguistic knowledge - in particular the human-likeness of their internal lexicon - remains uncertai…
Text GenerationEstimation de la qualit\'e d'un syst\`eme de reconnaissance de la parole pour une t\^ache de compr\'ehension (Quality estimation of a Speech Recognition System for a Spoken Language Understanding task)
Nous nous int{\'e}ressons {\`a} l{'}{\'e}valuation de la qualit{\'e} des syst{\`e}mes de reconnaissance de la parole {\'e}tant donn{\'e} une t{\^a}che de compr{\'e}hension. L{'}objectif de ce travail est de fournir un ou…
speech-recognitionSpeech RecognitionSpoken Language Understanding