paper-with-me

홈 › Papers

BAID: A Benchmark for Bias Assessment of AI Detectors

2025-12-12 · Priyam Basu, Yunfeng Zhang, Vipul Raheja arxiv

AI-generated text detectors have recently gained adoption in educational and professional contexts. Prior research has uncovered isolated cases of bias, particularly against English Language Learners (ELLs) however, there is a lack of systematic evaluation of such systems across broader sociolinguistic factors. In this work, we propose BAID, a comprehensive evaluation framework for AI detectors across various types of biases. As a part of the framework, we introduce over 200k samples spanning 7 major categories: demographics, age, educational grade level, dialect, formality, political leaning, and topic. We also generated synthetic versions of each sample with carefully crafted prompts to preserve the original content while reflecting subgroup-specific writing styles. Using this, we evaluate four open-source state-of-the-art AI text detectors and find consistent disparities in detection performance, particularly low recall rates for texts from underrepresented groups. Our contributions provide a scalable, transparent approach for auditing AI detectors and emphasize the need for bias-aware evaluation before these tools are deployed for public use.

📄 PDF Abstract BibTeX arXiv:2512.11505

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Large Scale Search Dataset for Unbiased Learning to Rank

2022-07-07 · Lixin Zou, Haitao Mao, Xiaokai Chu, Jiliang Tang 외

The unbiased learning to rank (ULTR) problem has been greatly advanced by recent deep learning techniques and well-designed debias algorithms. However, promising results on the existing benchmark datasets may not be exte…

Causal DiscoveryLanguage ModellingLearning-To-RankMeta-Learning+1

Unbiased Learning to Rank Meets Reality: Lessons from Baidu's Large-Scale Search Dataset

2024-04-03 · Philipp Hager, Romain Deffayet, Jean-Michel Renders, Onno Zoeter 외

Unbiased learning-to-rank (ULTR) is a well-established framework for learning from user clicks, which are often biased by the ranker collecting the data. While theoretically justified and extensively tested in simulation…

Learning-To-Rank

Towards Artistic Image Aesthetics Assessment: a Large-scale Dataset and a New Method

2023-03-27 · CVPR 2023 1 · Ran Yi, Haoyuan Tian, Zhihao Gu, Yu-Kun Lai 외

Image aesthetics assessment (IAA) is a challenging task due to its highly subjective nature. Most of the current studies rely on large-scale datasets (e.g., AVA and AADB) to learn a general model for all kinds of photogr…

Adversarial Examples Versus Cloud-based Detectors: A Black-box Empirical Study

2019-01-04 · Xurong Li, Shouling Ji, Meng Han, Juntao Ji 외

Deep learning has been broadly leveraged by major cloud providers, such as Google, AWS and Baidu, to offer various computer vision related services including image classification, object identification, illegal image det…

General Classificationimage-classificationImage ClassificationPornography Detection+1

Understanding the Effects of the Baidu-ULTR Logging Policy on Two-Tower Models

2024-09-18 · Morris de Haan, Philipp Hager

Despite the popularity of the two-tower model for unbiased learning to rank (ULTR) tasks, recent work suggests that it suffers from a major limitation that could lead to its collapse in industry applications: the problem…

Learning-To-Rank