paper-with-me

Papers

A Comprehensive Study on NLP Data Augmentation for Hate Speech Detection: Legacy Methods, BERT, and LLMs

2024-03-30 · Md Saroar Jahan, Mourad Oussalah, Djamila Romaissa Beddia, Jhuma Kabir Mim, Nabil Arhab

The surge of interest in data augmentation within the realm of NLP has been driven by the need to address challenges posed by hate speech domains, the dynamic nature of social media vocabulary, and the demands for large-scale neural networks requiring extensive training data. However, the prevalent use of lexical substitution in data augmentation has raised concerns, as it may inadvertently alter the intended meaning, thereby impacting the efficacy of supervised machine learning models. In pursuit of suitable data augmentation methods, this study explores both established legacy approaches and contemporary practices such as Large Language Models (LLM), including GPT in Hate Speech detection. Additionally, we propose an optimized utilization of BERT-based encoder models with contextual cosine similarity filtration, exposing significant limitations in prior synonym substitution methods. Our comparative analysis encompasses five popular augmentation techniques: WordNet and Fast-Text synonym replacement, Back-translation, BERT-mask contextual augmentation, and LLM. Our analysis across five benchmarked datasets revealed that while traditional methods like back-translation show low label alteration rates (0.3-1.5%), and BERT-based contextual synonym replacement offers sentence diversity but at the cost of higher label alteration rates (over 6%). Our proposed BERT-based contextual cosine similarity filtration markedly reduced label alteration to just 0.05%, demonstrating its efficacy in 0.7% higher F1 performance. However, augmenting data with GPT-3 not only avoided overfitting with up to sevenfold data increase but also improved embedding space coverage by 15% and classification F1 score by 1.4% over traditional methods, and by 0.8% over our method.

📄 PDF Abstract BibTeX arXiv:2404.00303

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationHate Speech DetectionTranslation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Hate Speech Detection using Large Language Models with Data Augmentation and Feature Enhancement

2026-03-05 · Brian Jing Hong Nge, Stefan Su, Thanh Thi Nguyen, Campbell Wilson 외 arxiv

This paper evaluates data augmentation and feature enhancement techniques for hate speech detection, comparing traditional classifiers, e.g., Delta Term Frequency-Inverse Document Frequency (Delta TF-IDF), with transform…

Hate Speech DetectionData Augmentation

Unsupervised Domain Adaptation for Hate Speech Detection Using a Data Augmentation Approach

2021-07-27 · Sheikh Muhammad Sarwar, Vanessa Murdock

Online harassment in the form of hate speech has been on the rise in recent years. Addressing the issue requires a combination of content moderation by people, aided by automatic detection methods. As content moderation …

Cultural Vocal Bursts Intensity PredictionData AugmentationDomain AdaptationHate Speech Detection+1

Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets

2024-07-02 · Kheir Eddine Daouadi, Yaakoub Boualleg, Kheir Eddine Haouaouchi

Today, hate speech classification from Arabic tweets has drawn the attention of several researchers. Many systems and techniques have been developed to resolve this classification task. Nevertheless, two of the major cha…

Data AugmentationEnsemble LearningHate Speech Detection

Hate Speech Detection in Turkish and Arabic: A Comprehensive Study

2026-06-30 · Somaiyeh Dehghan, Gökçe Uludoğan, Mehmet Umut Şen, Elif Erol 외 arxiv

Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate…

Hate Speech Detection

Enhancing social network hate detection using back translation and GPT-3 augmentations during training and test-time

2023-06-17 · Information Fusion 2023 6 · Seffi Cohen, Dan Presil, Or Katz, Ofir Arbili 외

Social media platforms have become an essential means of communication, but they also serve as a breeding ground for hateful content. Detecting hate speech accurately is challenging due to factors such as slang and impli…

Hate Speech Detection