paper-with-me

홈 › Papers

Neural Models for Offensive Language Detection

2021-05-30 · Ehab Hamdy

Offensive language detection is an ever-growing natural language processing (NLP) application. This growth is mainly because of the widespread usage of social networks, which becomes a mainstream channel for people to communicate, work, and enjoy entertainment content. Many incidents of sharing aggressive and offensive content negatively impacted society to a great extend. We believe contributing to improving and comparing different machine learning models to fight such harmful contents is an important and challenging goal for this thesis. We targeted the problem of offensive language detection for building efficient automated models for offensive language detection. With the recent advancements of NLP models, specifically, the Transformer model, which tackled many shortcomings of the standard seq-to-seq techniques. The BERT model has shown state-of-the-art results on many NLP tasks. Although the literature still exploring the reasons for the BERT achievements in the NLP field. Other efficient variants have been developed to improve upon the standard BERT, such as RoBERTa and ALBERT. Moreover, due to the multilingual nature of text on social media that could affect the model decision on a given tween, it is becoming essential to examine multilingual models such as XLM-RoBERTa trained on 100 languages and how did it compare to unilingual models. The RoBERTa based model proved to be the most capable model and achieved the highest F1 score for the tasks. Another critical aspect of a well-rounded offensive language detection system is the speed at which a model can be trained and make inferences. In that respect, we have considered the model run-time and fine-tuned the very efficient implementation of FastText called BlazingText that achieved good results, which is much faster than BERT-based models.

📄 PDF Abstract BibTeX arXiv:2106.14609

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Weight Decay 설명 없음
WordPiece 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Cross-Cultural Transfer Learning for Chinese Offensive Language Detection

2023-03-31 · Li Zhou, Laura Cabello, Yong Cao, Daniel Hershcovich

Detecting offensive language is a challenging task. Generalizing across different cultures and languages becomes even more challenging: besides lexical, syntactic and semantic differences, pragmatic aspects such as cultu…

Cultural Vocal Bursts Intensity PredictionFew-Shot LearningTransfer Learning

COLD: A Benchmark for Chinese Offensive Language Detection

2022-01-16 · Jiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 외

Offensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models. However, this task in Chinese is still under exploration due to the scarci…

A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages

2024-06-04 · Saminu Mohammad Aliyu, Gregory Maksha Wajiga, Muhammad Murtala

The proliferation of online offensive language necessitates the development of effective detection mechanisms, especially in multilingual contexts. This study addresses the challenge by developing and introducing novel d…

Hate Speech Detection

KOLD: Korean Offensive Language Dataset

2022-05-23 · Younghoon Jeong, Juhyun Oh, Jaimeen Ahn, Jongwon Lee 외

Recent directions for offensive language detection are hierarchical modeling, identifying the type and the target of offensive language, and interpretability with offensive span annotation and prediction. These improveme…

ArticlesClassification

Offensive language detection in Hebrew: can other languages help?

2022-06-01 · LREC 2022 6 · Marina Litvak, Natalia Vanetik, Chaya Liebeskind, Omar Hmdia 외

Unfortunately, offensive language in social media is a common phenomenon nowadays. It harms many people and vulnerable groups. Therefore, automated detection of offensive language is in high demand and it is a serious ch…