paper-with-me

홈 › Papers

A Hierarchically-Labeled Portuguese Hate Speech Dataset

2019-08-01 · WS 2019 8 · Paula Fortuna, Jo{\~a}o Rocha da Silva, Juan Soler-Company, Leo Wanner, S{\'e}rgio Nunes

Over the past years, the amount of online offensive speech has been growing steadily. To successfully cope with it, machine learning are applied. However, ML-based techniques require sufficiently large annotated datasets. In the last years, different datasets were published, mainly for English. In this paper, we present a new dataset for Portuguese, which has not been in focus so far. The dataset is composed of 5,668 tweets. For its annotation, we defined two different schemes used by annotators with different levels of expertise. Firstly, non-experts annotated the tweets with binary labels ({}hate{'} vs. {}no-hate{'}). Secondly, expert annotators classified the tweets following a fine-grained hierarchical multiple label scheme with 81 hate speech categories in total. The inter-annotator agreement varied from category to category, which reflects the insight that some types of hate speech are more subtle than others and that their detection depends on personal perception. This hierarchical annotation scheme is the main contribution of the presented work, as it facilitates the identification of different types of hate speech and their intersections. To demonstrate the usefulness of our dataset, we carried a baseline classification experiment with pre-trained word embeddings and LSTM on the binary classified data, with a state-of-the-art outcome.

📄 PDF Abstract BibTeX

Code (1)

paulafortuna/Portuguese-Hate-Speech-Dataset 공식 구현

Tasks

Word Embeddings

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

TuPy-E: detecting hate speech in Brazilian Portuguese social media with a novel dataset and comprehensive analysis of models

2023-12-29 · Felipe Oliveira, Victoria Reis, Nelson Ebecken

Social media has become integral to human interaction, providing a platform for communication and expression. However, the rise of hate speech on these platforms poses significant risks to individuals and communities. De…

Hate Speech Detection

ToxSyn-PT: A Large-Scale Synthetic Dataset for Hate Speech Detection in Portuguese

2025-06-11 · Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Diogo Fernandes Costa Silva 외

We present ToxSyn-PT, the first large-scale Portuguese corpus that enables fine-grained hate-speech classification across nine legally protected minority groups. The dataset contains 53,274 synthetic sentences equally di…

Hate Speech DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection

2021-03-27 · LREC 2022 6 · Francielle Alves Vargas, Isabelle Carvalho, Fabiana Rodrigues de Góes, Fabrício Benevenuto 외

Due to the severity of the social media offensive and hateful comments in Brazil, and the lack of research in Portuguese, this paper provides the first large-scale expert annotated corpus of Brazilian Instagram comments …

BIG-bench Machine LearningBinary ClassificationHate Speech Detection

Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages

2025-10-13 · Paloma Piot, José Ramom Pichel Campos, Javier Parapar arxiv

Hate speech poses a serious threat to social cohesion and individual well-being, particularly on social media, where it spreads rapidly. While research on hate speech detection has progressed, it remains largely focused …

Hate Speech Detection

Deep Learning Models for Multilingual Hate Speech Detection

2020-04-14 · Sai Saketh Aluru, Binny Mathew, Punyajoy Saha, Animesh Mukherjee

Hate speech detection is a challenging problem with most of the datasets available in only one language: English. In this paper, we conduct a large scale analysis of multilingual hate speech in 9 languages from 16 differ…

Deep LearningHate Speech DetectionQuestion Similarityzero-shot-classification+1