paper-with-me

홈 › Papers

Korean Online Hate Speech Dataset for Multilabel Classification: How Can Social Science Improve Dataset on Hate Speech?

2022-04-07 · TaeYoung Kang, Eunrang Kwon, Junbum Lee, Youngeun Nam, Junmo Song, JeongKyu Suh

We suggest a multilabel Korean online hate speech dataset that covers seven categories of hate speech: (1) Race and Nationality, (2) Religion, (3) Regionalism, (4) Ageism, (5) Misogyny, (6) Sexual Minorities, and (7) Male. Our 35K dataset consists of 24K online comments with Krippendorff's Alpha label accordance of .713, 2.2K neutral sentences from Wikipedia, 1.7K additionally labeled sentences generated by the Human-in-the-Loop procedure and rule-generated 7.1K neutral sentences. The base model with 24K initial dataset achieved the accuracy of LRAP .892, but improved to .919 after being combined with 11K additional data. Unlike the conventional binary hate and non-hate dichotomy approach, we designed a dataset considering both the cultural and linguistic context to overcome the limitations of western culture-based English texts. Thus, this paper is not only limited to presenting a local hate speech dataset but extends as a manual for building a more generalized hate speech dataset with diverse cultural backgrounds based on social science perspectives.

📄 PDF Abstract BibTeX arXiv:2204.03262

Code (1)

sgunderscore/hatescore-korean-hate-speech 공식 구현

Tasks

Cultural Vocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

How Does the Hate Speech Corpus Concern Sociolinguistic Discussions? A Case Study on Korean Online News Comments

2021-12-01 · NLP4DH (ICON) 2021 12 · Won Ik Cho, Jihyung Moon

Social consensus has been established on the severity of online hate speech since it not only causes mental harm to the target, but also gives displeasure to the people who read it. For Korean, the definition and scope o…

K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment

2022-08-23 · COLING 2022 10 · Jean Lee, Taejun Lim, Heejun Lee, Bogeun Jo 외

Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate…

Hate Speech DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

BEEP! Korean Corpus of Online News Comments for Toxic Speech Detection

2020-05-26 · WS 2020 7 · Jihyung Moon, Won Ik Cho, Junbum Lee

Toxic comments in online platforms are an unavoidable social issue under the cloak of anonymity. Hate speech detection has been actively done for languages such as English, German, or Italian, where manually labeled corp…

Hate Speech Detection

K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific Ratings

2023-10-24 · Chaewon Park, Soohwan Kim, Kyubyong Park, Kunwoo Park

Numerous datasets have been proposed to combat the spread of online hate. Despite these efforts, a majority of these resources are English-centric, primarily focusing on overt forms of hate. This research gap calls for d…

Hate Speech Detection

KoMultiText: Large-Scale Korean Text Dataset for Classifying Biased Speech in Real-World Online Services

2023-10-06 · Dasol Choi, Jooyoung Song, Eunsun Lee, JinWoo Seo 외

With the growth of online services, the need for advanced text classification algorithms, such as sentiment analysis and biased text detection, has become increasingly evident. The anonymous nature of online services oft…

Hate Speech DetectionMulti-Task LearningSentiment Analysistext-classification+2