paper-with-me

홈 › Papers

ToxSyn-PT: A Large-Scale Synthetic Dataset for Hate Speech Detection in Portuguese

2025-06-11 · Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Diogo Fernandes Costa Silva, Arlindo Rodrigues Galvão Filho

We present ToxSyn-PT, the first large-scale Portuguese corpus that enables fine-grained hate-speech classification across nine legally protected minority groups. The dataset contains 53,274 synthetic sentences equally distributed between minorities groups and toxicity labels. ToxSyn-PT is created through a novel four-stage pipeline: (1) a compact, manually curated seed; (2) few-shot expansion with an instruction-tuned LLM; (3) paraphrase-based augmentation; and (4) enrichment, plus additional neutral texts to curb overfitting to group-specific cues. The resulting corpus is class-balanced, stylistically diverse, and free from the social-media domain that dominate existing Portuguese datasets. Despite domain differences with traditional benchmarks, experiments on both binary and multi-label classification on the corpus yields strong results across five public Portuguese hate-speech datasets, demonstrating robust generalization even across domain boundaries. The dataset is publicly released to advance research on synthetic data and hate-speech detection in low-resource settings.

📄 PDF Abstract BibTeX arXiv:2506.10245

Code (0)

등록된 구현이 없습니다.

Tasks

Hate Speech DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Similar Papers 제목 키워드 기반

Large-Scale Hate Speech Detection with Cross-Domain Transfer

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Hate speech towards people with different backgrounds is a major problem observed in social media. Although there are various attempts to detect hate speech automatically via supervised learning models, the performance o…

Hate Speech Detection

Large-Scale Hate Speech Detection with Cross-Domain Transfer

2022-03-02 · LREC 2022 6 · Cagri Toraman, Furkan Şahinuç, Eyup Halit Yilmaz

The performance of hate speech detection models relies on the datasets on which the models are trained. Existing datasets are mostly prepared with a limited number of instances or hate domains that define hate topics. Th…

Hate Speech DetectionTransfer Learning

Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech

2021-09-01 · Findings (EMNLP) 2021 11 · Tomer Wullach, Amir Adler, Einat Minkov

Automatic hate speech detection is hampered by the scarcity of labeled datasetd, leading to poor generalization. We employ pretrained language models (LMs) to alleviate this data bottleneck. We utilize the GPT LM for gen…

Hate Speech Detection

ChatEMG: Synthetic Data Generation to Control a Robotic Hand Orthosis for Stroke

2024-06-17 · Jingxi Xu, Runsheng Wang, Siqi Shang, Ava Chen 외

Intent inferral on a hand orthosis for stroke patients is challenging due to the difficulty of data collection. Additionally, EMG signals exhibit significant variations across different conditions, sessions, and subjects…

Synthetic Data Generation

Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations

2026-03-18 · Dang H. Dang, Jelena Mitrovi, Michael Granitzer arxiv

We study whether large-scale unlabelled web data and LLM-based synthetic annotations can improve multilingual hate speech detection. Starting from texts crawled via OpenWebSearch.eu~(OWS) in four languages (English, Germ…

Hate Speech DetectionLanguage Modelling