Multilingual offensive lexicon annotated with contextual information
Online hate speech and offensive comments detection is not a trivial research problem since pragmatic (contextual) factors influence what is considered offensive. Moreover, offensive terms are hardly found in classical lexical resources such as wordnets, sentiment, and emotion lexicons. In this paper, we embrace the challenges and opportunities of the area and introduce the first multilingual offensive lexicon (MOL), which is composed of 1,000 explicit and implicit pejorative terms and expressions annotated with contextual information. The terms and expressions were manually extracted by a specialist from Instagram abusive comments originally written in Portuguese and manually translated by American English, Latin American Spanish, African French, and German native speakers. Each expression was annotated by three different annotators, producing high human inter-annotator agreement. Accordingly, this resource provides a new perspective to explore abusive language detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Abusive LanguageMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Quality Type-aware Annotated Corpus and Lexicon for Harassment Research
Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks.…
Vocal Bursts Type PredictionIdentifying Offensive Expressions of Opinion in Context
Classic information extraction techniques consist in building questions and answers about the facts. Indeed, it is still a challenge to subjective information extraction systems to identify opinions and feelings in conte…
Contextual-Lexicon Approach for Abusive Language Detection
Since a lexicon-based approach is more elegant scientifically, explaining the solution components and being easier to generalize to other applications, this paper provides a new approach for offensive language and hate s…
Abusive LanguageHate Speech DetectionMultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
Hate speech is a pressing issue in modern society, with significant effects both online and offline. Recent research in hate speech detection has primarily centered on text-based media, largely overlooking multimodal con…
Hate Speech DetectionVideo ClassificationPin\_cod\_ at SemEval-2020 Task 12: Injecting Lexicons into Bidirectional Long Short-Term Memory Networks to Detect Turkish Offensive Tweets
This paper describes a system (pin{\_}cod{\_}) built for SemEval 2020 Task 12: OffensEval: Multilingual Offensive Language Identification in Social Media (Zampieri et al., 2020). I present the system based on the archite…
Language Identification