paper-with-me

Papers

Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

2020-06-30 · NeurIPS 2020 12 · Robert Geirhos, Kristof Meding, Felix A. Wichmann

A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) use the same strategy. Accuracy alone cannot distinguish between strategies: two systems may achieve similar accuracy with very different strategies. The need to differentiate beyond accuracy is particularly pressing if two systems are near ceiling performance, like Convolutional Neural Networks (CNNs) and humans on visual object recognition. Here we introduce trial-by-trial error consistency, a quantitative analysis for measuring whether two decision making systems systematically make errors on the same inputs. Making consistent errors on a trial-by-trial basis is a necessary condition for similar processing strategies between decision makers. Our analysis is applicable to compare algorithms with algorithms, humans with humans, and algorithms with humans. When applying error consistency to object recognition we obtain three main findings: (1.) Irrespective of architecture, CNNs are remarkably consistent with one another. (2.) The consistency between CNNs and human observers, however, is little above what can be expected by chance alone -- indicating that humans and CNNs are likely implementing very different strategies. (3.) CORnet-S, a recurrent model termed the "current best model of the primate ventral visual stream", fails to capture essential characteristics of human behavioural data and behaves essentially like a standard purely feedforward ResNet-50 in our analysis. Taken together, error consistency analysis suggests that the strategies used by human and machine vision are still very different -- but we envision our general-purpose error consistency analysis to serve as a fruitful tool for quantifying future progress.

📄 PDF Abstract BibTeX arXiv:2006.16736

Code (1)

wichmann-lab/error-consistency 공식 구현

Tasks

Decision MakingObject Recognition

Similar Papers 제목 키워드 기반

Vehicle Fuel Optimization Under Real-World Driving Conditions: An Explainable Artificial Intelligence Approach

2021-07-13 · Alberto Barbado, Óscar Corcho

Fuel optimization of diesel and petrol vehicles within industrial fleets is critical for mitigating costs and reducing emissions. This objective is achievable by acting on fuel-related factors, such as the driving behavi…

Explainable artificial intelligence

Conformal Prediction Sets Improve Human Decision Making

2024-01-24 · Jesse C. Cresswell, Yi Sui, Bhargava Kumar, Noël Vouitsis

In response to everyday queries, humans explicitly signal uncertainty and offer alternative answers when they are unsure. Machine learning models that output calibrated prediction sets through conformal prediction mimic …

Conformal PredictionDecision MakingPrediction

Explainable AI to Improve Machine Learning Reliability for Industrial Cyber-Physical Systems

2026-01-22 · Annemarie Jutte, Uraz Odyurt arxiv

Industrial Cyber-Physical Systems (CPS) are sensitive infrastructure from both safety and economics perspectives, making their reliability critically important. Machine Learning (ML), specifically deep learning, is incre…

JAS-GAN: Generative Adversarial Network Based Joint Atrium and Scar Segmentations on Unbalanced Atrial Targets

2021-05-01 · Jun Chen, Guang Yang, Habib Khan, Heye Zhang 외

Automated and accurate segmentations of left atrium (LA) and atrial scars from late gadolinium-enhanced cardiac magnetic resonance (LGE CMR) images are in high demand for quantifying atrial scars. The previous quantifica…

Generative Adversarial NetworkSegmentation

The Epi-LLM Framework: probing LLM behavioral priors through epidemiological agent-based models

2026-06-01 · Petra Ferencz, Ava Keeling, Tobias O'Keefe, Lorenzo Stigliano 외 arxiv

Human behaviour during epidemics affects infectious disease dynamics, but quantifying this remains deeply challenging. Here we introduce the Epi-LLM framework: a novel integration of agent-based modelling, real-life epig…