paper-with-me

홈 › Papers

Independent Ethical Assessment of Text Classification Models: A Hate Speech Detection Case Study

2021-07-19 · Amitoj Singh, Jingshu Chen, Lihao Zhang, Amin Rasekh, Ilana Golbin, Anand Rao

An independent ethical assessment of an artificial intelligence system is an impartial examination of the system's development, deployment, and use in alignment with ethical values. System-level qualitative frameworks that describe high-level requirements and component-level quantitative metrics that measure individual ethical dimensions have been developed over the past few years. However, there exists a gap between the two, which hinders the execution of independent ethical assessments in practice. This study bridges this gap and designs a holistic independent ethical assessment process for a text classification model with a special focus on the task of hate speech detection. The assessment is further augmented with protected attributes mining and counterfactual-based analysis to enhance bias assessment. It covers assessments of technical performance, data bias, embedding bias, classification bias, and interpretability. The proposed process is demonstrated through an assessment of a deep hate speech detection model.

📄 PDF Abstract BibTeX arXiv:2108.07627

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualHate Speech Detectiontext-classificationText Classification

Similar Papers 제목 키워드 기반

Democratizing Ethical Assessment of Natural Language Generation Models

2022-06-30 · Amin Rasekh, Ian Eisenberg

Natural language generation models are computer systems that generate coherent language when prompted with a sequence of words as context. Despite their ubiquity and many beneficial applications, language generation mode…

Text Generation

Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection

2024-12-14 · Ahmed Haj Ahmed, Rui-Jie Yew, Xerxes Minocher, Suresh Venkatasubramanian

Social media platforms have become central to global communication, yet they also facilitate the spread of hate speech. For underrepresented dialects like Levantine Arabic, detecting hate speech presents unique cultural,…

Hate Speech Detection

CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network

2023-03-02 · Sreyan Ghosh, Manan Suri, Purva Chiniya, Utkarsh Tyagi 외

The tremendous growth of social media users interacting in online conversations has led to significant growth in hate speech, affecting people from various demographics. Most of the prior works focus on detecting explici…

Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection

2024-03-12 · Tharindu Kumarage, Amrita Bhattacharjee, Joshua Garland

Large language models (LLMs) excel in many diverse applications beyond language generation, e.g., translation, summarization, and sentiment analysis. One intriguing application is in text classification. This becomes per…

Hate Speech DetectionSentiment Analysistext-classificationText Classification+1

Prediction Uncertainty Estimation for Hate Speech Classification

2019-09-16 · Kristian Miok, Dong Nguyen-Doan, Blaž Škrlj, Daniela Zaharie 외

As a result of social network popularity, in recent years, hate speech phenomenon has significantly increased. Due to its harmful effect on minority groups as well as on large communities, there is a pressing need for ha…

Bayesian InferenceClassificationGeneral ClassificationHate Speech Detection+3