HateBR: Large expert annotated corpus of Brazilian Instagram comments for abusive language detection
Due to the severity of the social media abusive comments in Brazil, and the lack of research in Portuguese, this paper provides the first large-scale annotated corpus of Brazilian Instagram comments for hate speech and offensive language detection on the web and social media. The HateBR corpus was collected from Brazilian Instagram comments of political personalities and manually annotated by specialists, being composed of 7,000 documents annotated according to three different layers: a binary classification (offensive versus non-offensive comments), offense-level classes (highly, moderately, and slightly offensive messages), as well as nine hate speech targets (xenophobia, racism, homophobia, sexism, religious intolerance, partyism, apology to the dictatorship, antisemitism, and fatphobia). Each comment was annotated by three different annotators and achieved high inter-annotator agreement.
Code (0)
등록된 구현이 없습니다.
Tasks
Abusive LanguageBinary ClassificationSimilar Papers 제목 키워드 기반
HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection
Due to the severity of the social media offensive and hateful comments in Brazil, and the lack of research in Portuguese, this paper provides the first large-scale expert annotated corpus of Brazilian Instagram comments …
BIG-bench Machine LearningBinary ClassificationHate Speech DetectionPropbank-Br: a Brazilian Treebank annotated with semantic role labels
This paper reports the annotation of a Brazilian Portuguese Treebank with semantic role labels following Propbank guidelines. A different language and a different parser output impact the task and require some decisions …
Machine TranslationQuestion AnsweringSemantic Role LabelingSelf-Explaining Hate Speech Detection with Moral Rationales
Hate speech detection models rely on surface-level lexical features, increasing vulnerability to spurious correlations and limiting robustness, cultural contextualization, and interpretability. We propose Supervised Mora…
Hate Speech DetectionTowards a General Abstract Meaning Representation Corpus for Brazilian Portuguese
Abstract Meaning Representation (AMR) is a recent and prominent semantic representation with good acceptance and several applications in the Natural Language Processing area. For English, there is a large annotated corpu…
Abstract Meaning RepresentationEssay-BR: a Brazilian Corpus of Essays
Automatic Essay Scoring (AES) is defined as the computer technology that evaluates and scores the written essays, aiming to provide computational models to grade essays either automatically or with minimal human involvem…