paper-with-me

Papers

ICON$^2$: Reliably Benchmarking Predictive Inequity in Object Detection

2023-06-07 · Sruthi Sudhakar, Viraj Prabhu, Olga Russakovsky, Judy Hoffman

As computer vision systems are being increasingly deployed at scale in high-stakes applications like autonomous driving, concerns about social bias in these systems are rising. Analysis of fairness in real-world vision systems, such as object detection in driving scenes, has been limited to observing predictive inequity across attributes such as pedestrian skin tone, and lacks a consistent methodology to disentangle the role of confounding variables e.g. does my model perform worse for a certain skin tone, or are such scenes in my dataset more challenging due to occlusion and crowds? In this work, we introduce ICON$^2$, a framework for robustly answering this question. ICON$^2$ leverages prior knowledge on the deficiencies of object detection systems to identify performance discrepancies across sub-populations, compute correlations between these potential confounders and a given sensitive attribute, and control for the most likely confounders to obtain a more reliable estimate of model bias. Using our approach, we conduct an in-depth study on the performance of object detection with respect to income from the BDD100K driving dataset, revealing useful insights.

📄 PDF Abstract BibTeX arXiv:2306.04482

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeAutonomous DrivingBenchmarkingFairnessObjectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Predictive Inequity in Object Detection

2019-02-21 · Benjamin Wilson, Judy Hoffman, Jamie Morgenstern

In this work, we investigate whether state-of-the-art object detection systems have equitable predictive performance on pedestrians with different skin tones. This work is motivated by many recent examples of ML and visi…

Objectobject-detectionObject Detection

Benchmarking sentiment analysis methods for large-scale texts: A case for using continuum-scored words and word shift graphs

2015-12-02 · Andrew J. Reagan, Brian Tivnan, Jake Ryland Williams, Christopher M. Danforth 외

The emergence and global adoption of social media has rendered possible the real-time estimation of population-scale sentiment, bearing profound implications for our understanding of human behavior. Given the growing ass…

BenchmarkingSentiment Analysis

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

2026-05-18 · Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya 외 arxiv

Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities across knowledge retrieval, reasoning, code generation, and tool use. Th…

Code Generation

Benchmarking ChatGPT and DeepSeek in April 2025: A Novel Dual Perspective Sentiment Analysis Using Lexicon-Based and Deep Learning Approaches

2025-09-16 · Maryam Mahdi Alhusseini, Mohammad-Reza Feizi-Derakhshi arxiv

This study presents a novel dual-perspective approach to analyzing user reviews for ChatGPT and DeepSeek on the Google Play Store, integrating lexicon-based sentiment analysis (TextBlob) with deep learning classification…

Sentiment Analysis

Inequity aversion improves cooperation in intertemporal social dilemmas

2018-03-23 · NeurIPS 2018 12 · Edward Hughes, Joel Z. Leibo, Matthew G. Phillips, Karl Tuyls 외

Groups of humans are often able to find ways to cooperate with one another in complex, temporally extended social dilemmas. Models based on behavioral economics are only able to explain this phenomenon for unrealistic st…

Multi-agent Reinforcement LearningReinforcement Learning