paper-with-me

Papers

DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects

2026-04-07 · Jason Lucas, Matt Murtagh, Ali Al-Lawati, Uchendu Uchendu, Adaku Uchendu, Dongwon Lee arxiv

Harmful content detectors, particularly disinformation classifiers, are predominantly developed and evaluated on Standard American English (SAE), leaving their robustness to dialectal variation unexplored. We present DIA-HARM, the first benchmark for evaluating disinformation detection robustness across 50 English dialects spanning U.S., British, African, Caribbean, and Asia-Pacific varieties. Using Multi-VALUE's linguistically grounded transformations, we introduce D-CUBE (Dialectal Disinformation Detection Corpus), a core corpus component of DIA-HARM comprising 195K samples derived from established disinformation benchmarks. Our evaluation of 16 detection models reveals systematic vulnerabilities: human-written dialectal content degrades detection by 1.4-3.6% F1, while AI-generated content remains stable. Fine-tuned transformers substantially outperform zero-shot LLMs (96.6% vs. 78.3% best-case F1), with some models exhibiting catastrophic failures exceeding 33% degradation on mixed content. Cross-dialectal transfer analysis across 2,450 dialect pairs shows that multilingual models (mDeBERTa: 97.2% average F1) generalize effectively, while monolingual models like RoBERTa and XLM-RoBERTa fail on dialectal inputs. These findings demonstrate that current disinformation detectors may systematically disadvantage hundreds of millions of non-SAE speakers worldwide. We release the DIA-HARM benchmark, including the D-CUBE corpus (https://github.com/jsl5710/dia-harm), and evaluation tools (https://jsl5710.github.io/dia-harm).

📄 PDF Abstract BibTeX arXiv:2604.05318

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StopHC: A Harmful Content Detection and Mitigation Architecture for Social Media Platforms

2024-11-09 · Ciprian-Octavian Truică, Ana-Teodora Constantinescu, Elena-Simona Apostol

The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation archit…

HOD: A Benchmark Dataset for Harmful Object Detection

2023-10-08 · Eungyeom Ha, Heemook Kim, Sung Chul Hong, Dongbin Na

Recent multi-media data such as images and videos have been rapidly spread out on various online services such as social network services (SNS). With the explosive growth of online media services, the number of image con…

Objectobject-detectionObject Detection

ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark

2025-06-12 · Kangwei Liu, Siyuan Cheng, Bozhong Tian, Xiaozhuan Liang 외

Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content…

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

2026-04-18 · Huije Lee, Jisu Shin, Hoyun Song, Changgeon Ko 외 arxiv

Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To address these issues, we propose a framewor…

Improving Harmful Text Detection with Joint Retrieval and External Knowledge

2025-04-03 · Zidong Yu, Shuo Wang, Nan Jiang, Weiqiang Huang 외

Harmful text detection has become a crucial task in the development and deployment of large language models, especially as AI-generated content continues to expand across digital platforms. This study proposes a joint re…

Computational EfficiencyKnowledge GraphsRetrievalText Detection