paper-with-me

Papers

Exposing the Illusion of Fairness: Auditing Vulnerabilities to Distributional Manipulation Attacks

2025-07-28 · Valentin Lafargue, Adriana Laurindo Monteiro, Emmanuelle Claeys, Laurent Risser, Jean-Michel Loubes arxiv

The rapid deployment of AI systems in high-stakes domains, including those classified as high-risk under the The EU AI Act (Regulation (EU) 2024/1689), has intensified the need for reliable compliance auditing. For binary classifiers, regulatory risk assessment often relies on global fairness metrics such as the Disparate Impact ratio, widely used to evaluate potential discrimination. In typical auditing settings, the auditee provides a subset of its dataset to an auditor, while a supervisory authority may verify whether this subset is representative of the full underlying distribution. In this work, we investigate to what extent a malicious auditee can construct a fairness-compliant yet representative-looking sample from a non-compliant original distribution, thereby creating an illusion of fairness. We formalize this problem as a constrained distributional projection task and introduce mathematically grounded manipulation strategies based on entropic and optimal transport projections. These constructions characterize the minimal distributional shift required to satisfy fairness constraints. To counter such attacks, we formalize representativeness through distributional distance based statistical tests and systematically evaluate their ability to detect manipulated samples. Our analysis highlights the conditions under which fairness manipulation can remain statistically undetected and provides practical guidelines for strengthening supervisory verification. We validate our theoretical findings through experiments on standard tabular datasets for bias detection. Code is publicly available at https://github.com/ValentinLafargue/Inspection.

📄 PDF Abstract BibTeX arXiv:2507.20708

Code (0)

등록된 구현이 없습니다.

Tasks

Bias Detection

Similar Papers 제목 키워드 기반

From Fragile to Certified: Wasserstein Audits of Group Fairness Under Distribution Shift

2025-09-30 · Ahmad-Reza Ehyaei, Golnoosh Farnadi, Samira Samadi arxiv

Group-fairness metrics (e.g., equalized odds) can vary sharply across resamples and are especially brittle under distribution shift, undermining reliable audits. We propose a Wasserstein distributionally robust framework…

Evaluating Black-Box Vulnerabilities with Wasserstein-Constrained Data Perturbations

2026-03-16 · Adriana Laurindo Monteiro, Jean-Michel Loubes arxiv

The growing use of Machine Learning (ML) tools comes with critical challenges, such as limited model explainability. We propose a global explainability framework that leverages Optimal Transport and Distributionally Robu…

The Illusionist's Prompt: Exposing the Factual Vulnerabilities of Large Language Models with Linguistic Nuances

2025-04-01 · Yining Wang, Yuquan Wang, Xi Li, Mi Zhang 외

As Large Language Models (LLMs) continue to advance, they are increasingly relied upon as real-time sources of information by non-expert users. To ensure the factuality of the information they provide, much research has …

Hallucination

MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

2026-08-26 · Tianshi Wang, Jingsong Wang, Yafei Huang, Fengling Li 외 arxiv

Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple h…

Exposing the Illusion of Erasure in Knowledge Editing for LLMs

2026-06-22 · Advik Raj Basani, Anshuman Chhabra arxiv

Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an …

knowledge editing