paper-with-me

홈 › Papers

Towards a Comprehensive Taxonomy and Large-Scale Annotated Corpus for Online Slur Usage

2020-11-01 · EMNLP (ALW) 2020 11 · Jana Kurrek, Haji Mohammad Saleem, Derek Ruths

Abusive language classifiers have been shown to exhibit bias against women and racial minorities. Since these models are trained on data that is collected using keywords, they tend to exhibit a high sensitivity towards pejoratives. As a result, comments written by victims of abuse are frequently labelled as hateful, even if they discuss or reclaim slurs. Any attempt to address bias in keyword-based corpora requires a better understanding of pejorative language, as well as an equitable representation of targeted users in data collection. We make two main contributions to this end. First, we provide an annotation guide that outlines 4 main categories of online slur usage, which we further divide into a total of 12 sub-categories. Second, we present a publicly available corpus based on our taxonomy, with 39.8k human annotated comments extracted from Reddit. This corpus was annotated by a diverse cohort of coders, with Shannon equitability indices of 0.90, 0.92, and 0.87 across sexuality, ethnicity, and gender. Taken together, our taxonomy and corpus allow researchers to evaluate classifiers on a wider range of speech containing slurs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Abusive Language

Similar Papers 제목 키워드 기반

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

2025-10-10 · Bolun Sun, Charles Chang, Yuen Yuen Ang, Ruotong Mu 외 arxiv

We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives annotated with a five-color taxonomy of clear and ambiguous language catego…

Ubuntu-fr: A Large and Open Corpus for Multi-modal Analysis of Online Written Conversations

2016-05-01 · LREC 2016 5 · Hern, Nicolas ez, Soufian Salim, Elizaveta Loginova Clouet

We present a large, free, French corpus of online written conversations extracted from the Ubuntu platform{'}s forums, mailing lists and IRC channels. The corpus is meant to support multi-modality and diachronic studies …

ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks

2019-09-04 · Michihiro Yasunaga, Jungo Kasai, Rui Zhang, Alexander R. Fabbri 외

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to thes…

Scientific Document SummarizationText Summarization

Computational Argumentation Quality Assessment in Natural Language

2017-04-01 · EACL 2017 4 · Henning Wachsmuth, Nona Naderi, Yufang Hou, Yonatan Bilu 외

Research on computational argumentation faces the problem of how to automatically assess the quality of an argument or argumentation. While different quality dimensions have been approached in natural language processing…

LSCP: Enhanced Large Scale Colloquial Persian Language Understanding

2020-03-13 · LREC 2020 5 · Hadi Abdi Khojasteh, Ebrahim Ansari, Mahdi Bohlouli

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resou…

Translation